sfspark.org: Berthold Reinwald Spark/ML Q&A with Alexy Khrabrov
Recording: sfspark.org: Berthold Reinwald Spark/ML Q&A with Alexy Khrabrov
Hello everybody. I'm Alexey Khrabrov, the organizer of the self-spark and friends meetup. And here we are on location at IBM Spark Technology Center. And here we also have Berthold Reinwald. Uh who is the principal research staff member at IBM Research at Almaden. Correct. Uh welcome Berthold. Thanks for having me
and thanks for having us at the IBM location. You're welcome. Uh so uh so you'll be giving a talk on Apache SystemML. Correct. Uh and I think everybody's really excited. We've seen uh Spark used for machine learning. Uh we've seen IBM uh announcing huge investment in Spark at uh Spark Summit here in San Francisco last summer. So I just wonder uh what is going on at IBM? What's uh why did you guys choose, you know, to open source SystemML? Uh how you know, it came to be? So can I can you give us some context, you know, on your work and and how kind of fits the IBM IBM agenda? Sure
So we started off with SystemML actually a couple of years ago uh before Spark was the Spark that it is today. Uh we developed it actually on on Hadoop MapReduce because analytics, especially machine learning, was uh a good workload uh for cluster computing. Okay? And as we developed the system, then Spark grew grew and grew. And at some point in time uh our investment on uh Hadoop we kind of wanted to revisit and we came to a decision point saying, you know, either we develop our own distributed framework or we grab Spark. Yes. And we decided to grab Spark. Uh grab meaning rebase SystemML onto uh Spark. Okay? For many many reasons because Spark has a lot of advantages
It has now like Spark SQL, very good for data preparation and graph processing. It's very good for model building, for iterative machine learning algorithm. 80% of the machine learning algorithms are iterative algorithms and then scoring as well. So, it seemed like a it's a perfect platform for us. So, we took our SystemML, rebased it onto Spark and then, you know, 2015 uh Beth Smith, she went give give a keynote at the summit here in San Francisco and went on stage, announced the big STC um you know, start as well as you know, IBM's commitment to open source uh as well as open sourcing SystemML, you know, contributing. And IBM has a lot a long long history of uh contributing software to open source and SystemML was just right for it. Nice. So, uh so, it's kind of you guys basically uh started with Hadoop, started with distributed system, right? The kind of big data uh context and then saw Spark emerging and uh I mean, that's awesome, right? So, it's I think uh I see in the community basically people running Spark, they they want machine learning, right? And so, can you talk a little bit about what goes in the SystemML? Kind of how is it structured? Like uh if I'm a software engineer and I just want to take something which, you know, just works, in quotes, am I able to do this? Like how much of a uh you know, machine learning expert should I be? Uh how kind of uh enterprise ready is is SystemML? Yes
So, it is enterprise ready. Actually, uh before we open sourced it, we put it into IBM uh big data product called Big Insights. It's it's part of it and uh it's used there. So, it's enterprise ready. Um and you know, as a software engineer, you definitely can use it, okay? But the main focus of SystemML is really as a data scientist, okay? And we want to make the data scientist very productive without having to drop down to distributed computing, without having to drop down to a programming language. So, SystemML exposes a a language that allows that is meant for the data scientist to quickly develop custom machine learning algorithms. Mhm. Okay, and try many, many variants of algorithms Mhm
until they're happy with it. Okay. So, they can very quickly iterate over it. Interesting. comes with a a collection of uh pre-built algorithms that, you know, we put out as {quote} "examples" and uh people to use uh for free and uh customize them to their own purposes. Mhm. Interesting. So, uh can you give an example? For instance, let's say I want to build a recommender system
Mhm. Uh how can I do this in SystemML? Okay. So, you can use uh one of our existing algorithms. Actually, as a running example, we have uh uh alternative least square algorithm built in. We have uh several matrix factorization algorithms in there. So, you can the core machine learning algorithm is already there. Okay. But, I'm sure you have your own own nuts and bolts uh that you would like to add to it, so you can just use our DML scripts
DML is the name of our language that provide and modify it to your own purposes. Interesting. There's a higher-level language, so we are strong believer believers in uh declarative machine learning. That's what we want to call it. Declarative in the sense of you as a data scientist write down the algorithm, leave it up to the compiler, our cost-based optimization in SystemML that generates the right execution plan. Mhm. I, you know, uh optimality uh you know, if the system is perfect, Yes. optimal execution plan to to run from, you know, an Iris data set all the way up to uh you know, a multiplied Netflix kind of example
Mhm. Okay. Um so, uh that's interesting. So, um uh let's say um a lot of data scientists already know Python. So, I see kind of, you know, uh maybe 90% of uh data scientists which are educated today and kind of, you know, kind of boot camps, you know, uh kind of few month courses, maybe yearly courses like Python is the lingua franca of of data science. So, they So, So, how hard would it be for somebody who basically knows Python to use this DML language? Very good. So, several answers. Um When we started off with SystemML, we looked around
Well, we actually started off bottom up. Uh, we looked at MapReduce at first and realized it you know that is really assembler programming. No data scientist in the right mind will ever use it at that level. even software engineers. Exactly. And then we started to build up bottom up and then at some point we realized, well, we have all these operators, but nobody can really use those operators. You have to put the language on top of it. Then we were shopping around you know what language what syntax should we choose, okay? And we shopped around, interviewed data scientists, and so forth, and we decided to go with an R-like syntax
Okay. So, we have an R-like syntax for DML, okay? Uh Later on we also added another syntax variant to it that we call PyDML. It's Python-like syntax, okay? But it's still you know even our R-like syntax is a separate from R. Mhm. PyDML is still separate from Python. But their syntax like so it's on ramping for an R developer or a Python developer should be fairly easy. Mhm. However, if you come from let's say an R environment, okay? Or from a Python environment, okay? Then we have created APIs and I can talk about those in in the presentation as well Mhm
to allow from Python to call out to SystemML and run your algorithms, okay? The same thing is true from R. IBM has built a package that we call Big R that allows you to call from R you know from the comfort of your laptop or RStudio or Rattle or whatever tools you're comfortable with, connect to your cluster and invoke our DML algorithms. Okay. Uh, interesting. So, can you talk a little bit about the kind of the the the architecture of SystemML? So, is it is it written in Java itself? Is it written in Scala? You know, how does it interact with Spark? And kind of you know, if I'm a data engineer, I want to like pump a lot of data into it. Like how how does it interact with like all that kind of infrastructure, data infrastructure of Spark? Yes. So, SystemML it's 100% implemented in Java. Okay
And we used endless syntax for for our grammar. We our runtime operators, they go against the Spark core API. Okay. And at the top of it, okay, we have obviously a Scala interfaces, Java interfaces, and Python interfaces. Nice. I I mean, full disclosure, I also help organize the Scala meetup. So, I'm a big fan of Scala. And kind of you know, so kind of here my question would be so uh you know, I kind of I take this kind of very reduced view that Spark is a Scala DSL
Right. So, basically, Martin Odersky had this great talk at last year's Big Data Scala conference. You know, essentially saying that, you know, kind of this points you know, Spark is a Scala DSL and they essentially it's taking Scala collections and putting them on a cluster. If you're a Scala programmer, you know, you know, how you would work with collections. And it's also very declarative. You know, functional programming, basically, you say, you know, here is my collection. I want to map, I want to reduce, and you essentially have very intuitive operations to work on data, right? So, and that's a piece, you know, you always have to do kind of some ETL. So, so it makes it super easy for Scala people to to run ETL on on Spark on Scala collections
They can just take and do the same thing on Spark. So, let's say I'm a I'm a Scala engineer. I know very well how to, you know, do ETL in in in Scala. It's super easy, right? I can I can write one-liners. And now I can do the same thing on on Spark with the same kind of one-liners. So, and let's say now I want to do that. And and then I also want to run some machine learning. Yes
So, how easy is it to kind of, you know, use my Scala API knowledge and therefore Spark API knowledge, and kind of, you know, drop into SystemML, call into that, and then get back? How interoperable kind of this languages are in terms of data, you know, data collections? Yes. So, two answers. Today's answer is uh we have defined a Scala API uh which allows you to call from Scala SystemML. And it allows you to take uh you know, if you if your Spark SQL for instance produced a an RDD or a data frame, okay, you can take this data frame and directly and and you can, you know, uh cache it in memory, you can feed it into SystemML. Mhm. SystemML operates on that data frame, and, you know, depending what your machine learning algorithm does, it can produce uh data frames that we give back to to the to the Scala host, so to speak. Okay, and then you can feed it into your streaming, you can feed it into whatever interactive environment you have. Okay
Well, that's one answer, but that's a very coarse uh integration today. Yes, through data frames. Through data frames, but it's more like data data in and data out. So, that's very that probably covers 80% of all the cases uh but, you know, SystemML is open source now, and we defined Jira's actually where we want to take it to the next level and define a a much more fine-grained uh Scala DSL for SystemML. Okay. So, where you can have like uh you know, like a matrix multiplication Yes. kind of operation or what's a whole bunch of other unary operations on those uh collections of data sets, and directly express them in um in Scala, okay, in whatever environment you are in, and under the covers, you know, we we batch up uh through some ASTs, and eventually hand over the heavy lifting work uh to SystemML for execution. Interesting
So, it would be a much Right, today I would characterize it that we have a coarse-grained integration, but if people are interested in it, okay, a fine-grained integration is definitely feasible. Other systems have done it as well. So, maybe you're familiar with with Flink. They have a language in there. Maybe you're familiar with Mahout. They have Mahout Scala. These are bindings and things like that. So, it's all feasible and doable
Interesting. I think that's that's great that came up because, you know, I recently was at the conference called Typelevel, which is a gathering of, you know, functionally inclined folks in Scala. And so, they, you know, there is a whole set of of projects which basically want to put this formal and already have things in in Scala. So, I think it would be very useful to kind of get this Jira in front of Scala folks. And, you know, hopefully somebody will jump and see like how we connect this, you know, to Scala and come up with DSLs, right? Because that's what Scala is good at. Yes. So, that's that's great. So, like, you know, I'll, you know, I'll take this URL
If you guys are watching this, I'll put the link under There you go. the comment, right? So, so, we'll have the link and, you know, if you want to help, hopefully, you know, you'll be able to do to do this. That's great. I want to ask a little bit about kind of the unique position you're in being in IBM, right? Because IBM is is a country. Uh-huh. IBM has half a million employees all around the world, right? 400,000. All right. Still bigger than some countries
Uh and and, right? And so, I think the unique strength is that you have already lots of, you know, paying customers who have specific needs, right? So, you can take these products and put it in front of them very quickly and see what they use case. So, I wonder, you know, how this process works for you, how much of exposure SystemML got in industry, do you have this kind of feedback loop like, I, you know, I've given it to to clients and their data scientists, and do you see what they're doing with this? Yes. Yes, all of it. Okay. So, we give it, you know, because IBM is a is a country, okay, we have our internal customers, so we give it to other groups to try it out first before we put it put it in front of external customers, so to speak. We did that as well in the form of betas, continuous betas, or in the form of POCs. So, very early on actually we we we went off and did POCs with customers to try out, okay, in terms of usability, in terms of capabilities, in terms of now what are the use cases that you really need to address, okay? Because you can build whatever you want, okay, but if it doesn't match and fix a customer's problem, then it's a pretty much useless. Right
Right. Right. Uh so, so are these internal groups like folks who go to, you know, customers and solve their machine learning problems? Like can you talk a little bit like what what are the domains? Like what kind of problems do you see? So, um the the POCs and the domains they come from all different kinds of verticals. So, we had customers from insurance, customers from automotive, customers from finance. They come from all all different verticals, okay? And and all the problems are different, okay? They're all kind of difficult to deal with because first you have to kind of understand the lingo a little bit, okay? And you know, nobody comes with a proper problem definition to you, okay? You kind of have to dig into it and understand it. And then especially in the case of feature case of machine learning, the entire feature engineering until you arrive at a at a model building kind of thing where you really want to shine, okay? You have you have to spend a lot of upfront time first in order to get the you know Yes. So, so basically you you are kind of testing different domains and you kind of I hope you like you you see commonalities and kind of Yes, exactly. Okay
So, actually in in the talk subsequently, so I have actually produced one slide where we kind of enumerate the the you know, top six or so use cases, you know, that we got to hear from our customers. And I can happily take you guys through it. Okay. Well, that's that's exciting. So, I guess you know, we'll kind of wrap up with these questions about the future, right? So, it's super exciting to have kind of we have a diversity of ML options, right? So, and and kind of what do you see like there are multiple libraries, right? So, what do you see as strengths for SystemML and kind of you know, who would make kind of the best users, right? If somebody is wants to pick it out and try, you know, like there are small startups, there are like bigger companies and there are data scientists or data engineers and sometimes people wear both hats, right? Like there is somebody who can use to set up a data pipeline and they have to capture data from an API, put it into Kafka, you know, load it into Spark, do some analysis. So, who kind of what kind of setups works best and kind of where do you think SystemML makes most sense to try? So, SystemML, in my opinion, it's different from all the existing systems. You call them libraries. Mhm
We really want to differentiate that we don't want to be a library. Okay. We want to be a provide a language for the data scientist Okay. to quickly develop their own machine learning algorithms without having to worry about, okay, do I run it on MapReduce, do I run it on Scala, do I need to have a GPU, how large of a cluster do I need to have, you know, on my training set I try it out on on 10,000 rows today, but tomorrow in deployment I might have a billion rows. Right. My same algorithm hopefully still works. Okay. So, you don't have to worry
So, we are proud of, you know, this declarative machine learning approach that we took. Mhm. We give you a language, okay? And this language is used to express your own machine learning algorithm. So, Got it. So, Maybe one more piece of background, maybe, because we, you know, IBM is large, has a long lot of history. So, uh but in IBM Almaden Research actually this is where the database systems actually come from, and four decades ago, you know, the relational database systems were created and the SQL standard was created around it, okay? Which was a big, big event, okay? Because it enabled application developers without having to worry about database systems to develop their applications, okay? And toolings and so forth. I see. So SystemML is the SQL for for machine learning
That's how we like to think of it. This is great. Now, this is a great ambitious goal. But then I have to ask you, you know, now if I'm a data scientist who kind of put together TML scripts, right? How easy is it that then can I take this in production immediately? Or do I have to kind of generate underlying Scala, Java, or how can I productize that algorithm? So, we we built a compiler, although it's it's not a compiler in the strict sense. It's really like a a compiler interpreter kind of thing. So, we compile, optimize it. So, it's it's declarative machine learning. We we look at data characteristics
We look at the characteristics and based on and and the operations that you need to perform and and based on that information, we create an execution plan, okay? We don't generate Scala. We don't generate Java, but we have our own runtime operators, like for instance, we have our own matrix library in a distributed case and we have you know probably seven or eight different ways to do matrix multiplications, okay? So, depending on whether you have a a tall skinny or a short and wide or a square or or sparse or dense matrix, depending on that, we choose the right runtime operator that we have in our library in order to get the job done. Okay. So, basically, I should be able to write the script and put it into production immediately. Yes. Excellent. Now, this is great news. And okay, so then I'll probably kind of my last question will be what is the future direction and when can where can community help? We already mentioned Scala DSL and I'm I'm obviously as a fan of Scala DSLs, you know, I'm like all for it
Like where else do you see most kind of progress and where would you like the community to jump in, you know? So, that's a a lot of work left to do. Mhm. Um and we actually on the Jira server we put a roadmap up as well, but the rough categories are really uh consumer ability. Mhm. Okay, so that is always uh the biggest uh you know, making this hurdle of taking system ML and putting it in your environment and running it. Have the right stable APIs that Mhm. that is extremely important, okay? Um in terms of engine itself, we have we have a cost-based optimizer there. There's a cost model there that needs to be tuned and enhanced, you know
There's a lot of problems still unsolved in there, so So, this is like research problems, too. Research problems as well. If somebody like wants to do a PhD, there is a room. Go for it. Download it. Add it, you know, write a thesis about it. Yeah, it's good. Yeah, cool
We are we are for it, okay? The market is there. The market is definitely there. So, that's the second category. The third category I would say Yes, we put out about 20 plus algorithms, Mhm. clustering, classification, blah blah blah. Mhm. There's more, okay? Believe it or not. We should we can write in this language
In this language, okay? And contribute it back to to the to the open source. Okay. Okay. Well, that sounds like a fun thing. That's the fourth thing. Um tooling is always an issue, okay? Uh yes, we we we produce optimal execution plans, but you know, really kind of visualizing or monitoring how well the system performs in a in an executed in in a distributed environment. Uh so, tooling for it that helps the system ML developers or maybe it's a user even to some degree uh to explain what's really happening as you execute this multi matrix multiplication. That would be helpful as well
And And another thing that a lot of people always forget, uh I think it's time to do benchmarking uh for uh machine learning. Yes. Okay? And that is an extremely difficult problem. And it will be a fun thing to watch. Yes. Do Do guys have some benchmarks already? Uh we we run stuff, but no, we do not have a benchmark. Cool. So, it sounds like something community can contribute, right? Like we need the neutral ground for this, right? So, that sounds like a cool community project
So, cool. Well, thank you very very much. We're looking forward to your talk, and we'll get these links from you with the Jiras and the road map, right? And we'll post it under the meet up and under the video, and you guys should be able to find it. Awesome. Thanks.