sfspark.org: Alexy Khrabrov interviews Denis Kulgavin
Recording: sfspark.org: Alexy Khrabrov interviews Denis Kulgavin
welcome everybody this is SF spark meet up and where location here at nitro where I'm a chief scientist and at the same time the organizer of SF spark which is the new meetup here in San Francisco bout spark fling related technologies kafka acha and mill date is an example of a streaming framework and today were very happy to have with us Dennis Cole Gavin who is the CEO of min data and one of the creators of the data framework for streaming I'm very excited to have him present this framework for the first time in the meetup setting at estes park today so first welcome to nadra alexi thanks for having us so tell me a little bit about the trajectory which brought you to creation of I mean data yeah that's a great question you know originally we come from a consulting background and so we built a number of distributed systems for the financial services and add text spaces almost all of the systems had very low latency high throughput requirements and in a lot of cases in the past we had used Apache store and out of building a lot of solutions on top of patchy storm we saw quite a number of areas that we thought were right for improvement and so we ended up building the mint data stream processing platform as essentially an extension of division that storm had started to see essentially how we can improve people's efficiency with string processing so in the burgh repeated measures that you build various embedded systems which were running on trains and buildings and these things are sources of stream and data so i wonder how did this background effect kind of your career and in what ways do you see mandate a kind of result of this imbalance directory yeah it's funny you mentioned that so in the late 90s I worked at a company called a challan for five years and in those days it's today what we call IOT or the Internet of Things in those days we called it embedded systems I and we had a lot of the same challenges we had a very large number of devices they were always emitting advances and everything was very much event-driven and so today when we call things big data and sort of data moving at high-velocity we find ourselves solving similar problems just at a very different scale there's more data that's moving around and so a lot of the ideas for tooling and mechanisms to make people more efficient that we had back at echelon I've decided to kind of try to bring forward to the big data space and so that's where a lot of the ideas behind the mint data run time as well as the tooling around it kind of comes from mm-hmm so it's interesting to see the streaming space evolve but I think of storm was the kind of pioneering framework knowledge of the public spark was developed at about the same time so that it's part streaming which is a different model which is mini batching and now we have new hampton project link which is again through me first and second so where has been day to fall on the spectrum and basically how do you see this whole application space going forward yeah that's that's an important point so I think that it's important for the string processing platform to take several things into account so the first part is the string processing engine this is an engine that semantically at a very high level is similar to storm so how spouts it has bolts it has the ability to distribute these things on a cluster I'd give certain on messaging guarantees at least once at most once exactly once so I think all the engines in the one I've been data that we've built from scratch certainly embody this feature second functionality from there you then start to say well if I have a high-performance engine how do I start to take advantage of it because you know our feeling it meant data is that it's not sufficient just happen to have a high-performance runtime and you actually want to sort of make people more efficient because people end up contributing to large portion of the cost at a company and so if we can make people more efficient that's really where the importance is and so if metadata yes we've built a high-performance runtime we're happy to show it off but we've also built a set of tooling that allows you to manage the stream processing designs that you built out it allows you to manage the environments on which they run and on top of that it actually allows you to build real time applications for the first time in a point-and-click fashion and so these are no longer dashboards you know right now I think that there are a number of solutions out there that sort of help you build dashboards but dashboards allow you don't have a one-way conversation with your data they present some information and most you can have a filter or two to kind of drill in on it a little bit but it's very different from having a true rich real-time application that can interact with that data set and be hooked up to a high performance for on time at the same time and so I think that as time goes on mint data and other real-time streaming platform vendors will actually start to provide much more sophisticated tooling that essentially not only allows you to define a screen processing topology or design as we call it but also the ability to build real-time streaming applications the same way that bi vendors you know decade or so ago allowed you in the era of crystal reports to very quickly built form say think we're going to see the same kind of powerful rich tooling gets trapped the high performance run times that we're just seeing along today right the I think that now makes clear to me why you guys have this beautiful green which is new to me in in the case of streaming systems and generally back-end systems are not known for having wonderful goods and and I wondering how do you envision your customers are they developers are the analysts basically who is the target audience of this goo is and how do you see your product being used you know in industry yeah so that's a very good question so they're really three constituencies that we take into account the first is obviously the business user so this is the person that's going to be consuming the real-time application this is really the reason that as developers you know we exist in the first place we're there to really serve the needs of the business and so the first person that served as is the business user that's going to be consuming the application the second person is the business analyst or data scientist sort of it's a number of roles at different companies but these are the people that are going to use these powerful go is to construct these real-time applications to construct and define the environments in which they run and then in the case of meant data it's as simple as sharing a link it's just like a Google dog but now you're sharing a true real-time application that's hooked up to a string processing runtime that's the second constituency and the third one is operations we can't forget operations operations are there to make sure that the things that run today won't break tomorrow and we have to take them into account we have to give them the same permissions and controls and flexibility to define who can do what where and why with exact disability into the performance and would cause analysis performance issues without really leaving them behind so I think that in the future stream processing platform vendors really need to take all three people into account then user that needs to have a beautiful Gilly and they want this built very quickly than the analyst that can construct this quickly and operations person that can control and define who has access to it okay this is great I think it's really important that you're taking into account all these considerations and divorce is a very important component which is often overlooked right that that's a pretty crucial production so that's really great to hear that so I know them you mentioned that your system is built on components and it was really intriguing when you mentioned that you can actually use spark is one of these components and so you primarily develop in Java and last time you talked about presenting at the meetup we were scholarship here have you sugar hyper service call so literally fennel scholar you basically said you guys can do skull edema and I said that the only if you have a hundred percent complete scar them so I'm wondering first you know how is it you know that spark is a component because usually thought as kind of the V kind of the Nexus of all the other things so it's very interest to me that you can use it as a component and secondly what were your experiences developing in Scala for this presentation and what their thoughts on using scholar for kind of end-to-end data pipelines in which is the topic of the upcoming bday the scale conference in in August and you know basically the typical use case for Scotland did the world now is to write significant pieces of it including our cacophonous bar in Scala I wonder what are your experiences yeah so let me kind of answer that in two parts so first of all with spark only thinks spark is an amazing technology and actually I spoke at a meetup a couple weeks ago and I gave our developers a challenge I said can you integrate spark onto our platform in one day mmm and they were able to pull this off and we are able to present a spark running on the mid data platform it's all because we have a very straightforward componentized architecture and we think that spark is an amazing technology I think a lot of people are going to be familiar with it especially with IBM's entrance into the space i think you know they mentioned that over a million people will be eventually trained on spark and the different transformations and operators I think spark has this amazing technology called catalyst that essentially optimizes sequel queries to run against the data or essentially with the data frame API has a very sort of high-performance way of taking a logical query plan and physically executing it and so all that richness and goodness and can exist and does exist on meditative platform and it took literally just a day to put it on the platform and we think spark is here to stay we think it's going to be a very exciting way to do this in the case of the data with spark either it can run on the mint data JVM itself or we can connect to a faraway spark cluster it depends on if you're more in a sort of data exploration mode or if you're crunching lots of data you may off to call out to the spark cluster but both modes are literally at the flip of a switch in the mint a DUI and you can use spark at will and we think it's a great complementary to go technology so that's in terms of spark in terms of Scala it's actually it's an amazing which I recall sort of when I was first learning Scala it was when I was very fascinated with the Kestrel q the Twitter had built a number of years back mm-hmm and having built some components for the upcoming presentation today in scala i can say that it's a very expressive amazing language and we're excited to to sort of walk in this call community to the men tato platform it's really a first class citizen alongside with Java to develop components on the platform all right this is very exciting and I think we'll wrap up for additional questions to ask for you know startup founders where are you guys division yourselves in the year can they give us some predictions and definitely have it back then and we'll check it check it out yeah that's great question i think we'll see wider adoption of string processing platforms i think what you'll see in the future is that today a large number of real-time applications or what people think of as dashboards that are built by hand from scratch you know i think people will stop building them from scratch and i think they'll use more powerful tooling to construct those applications and i think we'll continue to use stream processing 1 times for many years to come oh ok and Mendez itself will hopefully be the poison multiple setups yeah so we have a number of customers quite a number of production we're looking to actually you know get this out to a wider audience but I think the important thing to really keep in mind is that it's all that human efficiency whether it's meant data or another vendor we really look forward to a future where the tools that are at our disposal make us more efficient overall all right this is a great prediction and we'll definitely will welcome you back even before that but certainly response the span of time and see how it goes out and we'll look authority at all great thank you so much thanks for having us here