Reactive Systems and Microservices: Michael Mahoney Interview
Recording: Reactive Systems and Microservices: Michael Mahoney Interview
hello everybody I'm Alexa crabber of the organizer of several meetups we should have adjournment up tonight we have reactive systems and we also have a soft spark and the campus of Purdue and the here also have Michael Mahoney who is on the faculty statistics at Berkeley and so Michael gave an amazing talk at dr. pamel about spark versus NP I wish kind of the question was I was always wondering mother Michael really kind of nailed it so we my team put together a talk here in the real glad to be here so I mean you know I kind of call the patient's won't have enough emails but maybe we'll touch on some so statistics is foundation of machine learning data science right some people say like that the census Buster died statistics and machine learning and AI rebranding of this right so what does you take on statistics how important it is to know statistics all right why don't we see a lot of statisticians basically the forefront of a guy right because statistics touched on all this large data sets Trane's right it's kind of a task all these missions but in we also always see them our package there brandon bastardized but we don't see this rigor where do you think this you know this happened and like how can we get more like statistical rigor into all this stuff yeah I mean I think that's a complicated question and a complicated problem I mean I think the flip side of what you're asking is why do you see more computer scientists at the forefront and la statistics and so to answer that question will answer the other question in computer science and statistics are very different culturally as well as technically and one of the sort of great strengths of I think computer science is the sort of can-do attitude we just go do stuff and it doesn't need to be sort of statistically right or rigorous you just do it and sometimes it works and sometimes it doesn't and so that's the great strength but the weakness of course also is that you sometimes just go do stuff and you don't have the statistical rigor the carefulness right and so in some applications there's a first to market effect and and so the you know the person who goes and does stuff initially is is the winner and and so in that sense maybe there's a selection bias or four people or research areas that do that and in other cases there's people statisticians or other people that are said a little bit more cautious and so it's certainly foundational but there's a real challenge I think to getting that level of rigor and care at scale for example which is more the domain of computer scientists is it so so one of the basically the problems is so immense that this decisions mostly live in academia right and academia is doing things their own way and kind of it's a kind of slower pace it's more peer-reviewed right and computer science is just going to do stuff right but the question is I mean how how does the t6 remain relevant there right because let's say you know the prosign is going run a bunch of stock on spark and they say here are some numbers but there is no confidence in these numbers right are they just basically give up so how important is it to kind of to to be sure that your numbers are correct do we need to go back to this do we need like you know there was this blink DB project like that you know error intervals and you know I heard about it then I don't hear about it anymore so do you think this is important for us to kind of when we get some number of summation learning to kind of understand or is it sufficient for you know for these numbers to work like give you some ads people click on the ads so they make enough money it doesn't matter anymore right like right they have a lot of clicks yeah I mean I think that depends strongly on the application area I mean one thing you said statistics is sort of it in academia and I think there's a stronger bite so so that's I think there's a lot of statisticians not in academia but sometimes they're in a biostatistics department or sometimes they're in econometrics department or sometimes they're in industry and so it's a little bit less cohesive or more diverse and so you see them in a lot of different areas hmm but sometimes they're they'll call themselves biologists or economists so you see methodology like this in a lot different areas and so I think I think that's one thing in terms of diversity now that doesn't address the second sort of question you had about is their value or if there is how do you integrate it into the sorts of computations that computer scientists regulate do and I think in some cases the sort of level of statistical rigor that is that decision might bring to bear is is less important mm-hmm in the sense that as I said if you're first to market and you have some result you know all you need to be is roughly right and or not but you know roughly right and then you can figure it out later now there's downsides to that and and the statistician would be more careful about managing those downsides but in certain types of environments such as ad clicking if you're roughly right that's good enough right on the other hand there's a lot of large-scale data that's beyond getting people to click on ads say personalized medicine and you're not gonna be very happy if you go to the doctor he's roughly right right and and you know you you're roughly cured but not completely correct so I think as large-scale data problems go beyond social media internet advertising to a range of other areas you will see and you are seeing now relative to ten or fifteen years ago and increasing sort of emphasis or desire for some of these statistical rigor questions so it sounds to me almost like this is very simple with the stereotype as it says using a small data and so R is the favorite tool of statisticians this main problem is not scalable right now so there is a whole bunch of companies we want to put things on are going spark right are on plaster right on distributed so so would you say that it's kind of you know good kind of role found that such decisions live in small data and just live in you know they can do big data yeah I don't know what small or big is and I was doing big data back before it was massive before as large so you have these words I think I mean certainly if you look at you know at things with a longer history statisticians are doing big data back in about oh one I'm like a 1801 yes back when they were trying to figure out the orbit of comets and planets and so on and developed you know the method of least squares say mmm-hmm so in the twentieth century there was I think a de-emphasis on computation within statistics as computer scientists and scientists computing embraced it more aggressively and so what you're seeing is a little bit of a manifestation of that so the so ours often times done at smaller scale that's true on the other hand are as a fabulous environment to do all sorts of computations in and that's something that computer scientists largely haven't had you know a nice environment so you have a MATLAB which talks to the numerical analysis in certain scientific communities now you're seeing Python and you know open source environment but that's in some ways similar to our and the ecosystem art has but in other ways very different and so there's a scale issue but you know there's a lot of things I mean modern computers this computer sitting over here probably has you know the version if you go out and buy it yesterday and get the largest memory has something on the order of a terabyte of data yes that's an enormous amount of data yeah there's you'd be hard-pressed to find problems where you really need to be larger than that and so that's pretty large small data if you're working on one machine yes used to be yeah this was big data ten years ago it was gigantic right more than the bigger biggest big data so so but I'm very curious so you basically got into spark which a lot of the decisions don't do so what prodded you right I would say I would posit that you're different from most decisions who stained are they may be using you know weather may be using spark a little bit but like you they really got into that like the the car questions of spark what probably you to take a look at spark yes I may be a peculiar statistician I'm a statistician has actually never taken a class of statistics and I also sit in the amp will have in computer science I'm a computer scientist has never taken a class in computer science all right so not knowing what you don't know and what you shouldn't do maybe it makes it a little bit easier to just go and do stuff and figure stuff out so we had been doing work in the theory of algorithms that had a very statistical flavor having to with randomized linear algebra randomized matrix algorithms and these algorithms you know algorithms in this class of albums and randomized linear algebra solved the only squares regression faster than Gaussian elimination so the best worst-case algorithm in existing cases algum randomized linear algebra if you want to do high-quality numerical computations in PDE applications partial differential equation applications algorithms from this area done in low-rank approximation x' i mean are the best algorithms in those area so we've done a lot of work on the theory sort of worst case computer science theory but also on understanding the statistical properties of the algorithms on large data by large I mean tenth of a terabytes it fits on one machine not obscenely large ok and at some point the question came you know how if at all is this stuff relevant for larger scale and there are the communication computation questions are just very different the types of problems you have a very different the questions you might want to ask are very different and so the question we had I got to know some people at Lawrence Berkeley National Lab and and and Cray and we decided to put together a project with motivate advice by some particular scientific questions we wanted to ask the question how would this class of atoms is randomized matrix algorithms perform on terabyte and up scale and what we quickly saw was that large scale means different things to different people so if you're a database person or computer scientist large scale means go to a distributed data center and do Hadoop or write a produced if you're scientific computer you're saying we'll go find a supercomputer and do high-quality numerical analysis and there's an enormous disconnect between those two areas people really talk past each other and fail to appreciate these subtleties and significance with the other side has to offer and so we tried to do is carve out a space there because when we tried to answer the question before doing that ourselves how do these algorithms perform in spark or in Hadoop on super terabytes scale data we couldn't get an answer people would tell us how they should perform but they were telling us sort of revealed more about the biases of their area than how this actually performed and so we went and did it it was a non-trivial amount of effort to actually figure out how this would perform but we asked the question for very particular motivating problems that weren't social media internet advertising because for there you don't need extremely high quality linear algebra basically because you just need to be roughly correct in your predictions yes but for a certain scientific applications you just needed much higher quality answers in you the stress test the linear algebra more and we wanted to ask the question if you wanted to do linear algebra such as these randomized matrix algorithms on super terabytes scale you could do it in MPI in a high-performance environment you could do it on a distributed data center and spark and or the trade offs but if you put spark here what if you and distributed data central strong it becomes what if you didn't use spark what did you lose by the convenience of having this and if you lose a factor of a thousand that may not be so good if you lose a factor but - that may that may be a price a lot of people are willing to pay even in high performance environments say you lose a factor of you know two to ten but then everything is geniuses everything is inspiring so you kind of win by not having to ship one to the other but you know I finally talked revealing and you know I'm super happy have season because they let the question ask also because I was exposed to high performance computing earlier and you know I was kind of spreading the knowledge of spark to bio people for us so I gave a talk at pencil bioinformatics and I asked you know who's using spark nobody raised their hand look at spark everybody raised their hand what I guess using no we use an MPI why because we were ought to be granted to an ACEF and we've got like a big cluster which is worth several million dollars which was a different grant cycle it was written five years ago so now look at this decoster everybody came to use in C++ and the question is you know children move to spark who can this be clustered obviously mp9 people will not kind of be in favor of this because they're ready you know order and so so that kind of scent sounded like it's kind of you know just conservative approach right but but obviously MPI people figured out something they would do it for years and this whole area of hpf is actually not very familiar to a lot of big data people in the valley because we took a dupe and they'll not be ready to spark and then we do typical things we should not watch the matrix competition until we come to deep learning right so so guys my next question is a lot of the deploring stuff and optimization is a mathematics algorithms and in the Milner algebra rather we do great at the sound so I mean I don't want to probably talk about do you think that it's now more relevant with the proliferation of deploring which eventually souvenir algebra to ask this question in this context so I think if you ask the question ten years ago where is their computationally intensive machine learning at scale you'd be hard-pressed to find it there's certain applications in some of the scientific examples our predecessors who would talk about that were computation Tencent as opposed to communication mm-hmm and you lose a factor of two to ten or whatever the numbers are in spark if you have computation if you have algorithms that have very straightforward communication patterns if you have algorithms that have very complicated communication patterns as most you know interesting linear algebra algorithms do you may lose a factor of twenty to a thousand so you potentially lose a lot more if you're not careful there's a question about whether or not those the computation do and run but but that's that's the case if you ask 10 years ago are those the computations you want to run you'd be a little bit more hard for us to find many examples so one perspective on the recent work and deep learning is that this is a nice clear sort of killer app that fits within the umbrella of machine learning where you're computationally intensive as opposed to communication intensive and one of the main bottlenecks is is good performance of the linear algebra yes a stochastic gradient descent in the details of the optimization are different than the the the matrix algorithms I'll be talking about today but but your computationally intensive in a way that you weren't typically in machine learning say ten years ago yes I remember see you you're basically do cover kind of the typical in or SVD right like a negative position but it's obviously cover the next question what we're gonna see and I mean I'm just curious because I think we're different is going is different apologists right they're basically reinventing more and more different ways and it just sounds like there will be even more communication patterns because they want to connect more layers and they want to give it attention and they want to give it memory right and it means the different layers will have to basically refer to this memory bank right so this sounds to me again I'm not a specialist than this I kind of I follow the field but it sounds to me like they'll be both interesting communication and both interesting reputation I think that's the case and I think it'll be very different than the computation communication trade-offs you see in linear algebra in scientific computing applications some of the the parameters you mentioned as well as others you know the size of a batch if you're doing a mini batch or the rate of dropout a range of other things you're deep learning is optimizing a sometimes it's called a free energy because minimizing an entropy cross entropy between you know the model and the data and so for that you need to do somewhat different computations but there's a range of parameters that you can trade off of each other and so the state-of-the-art deep learning has a lot of parameters that you play off of each other to get to the high performance but if you take a step back ants what are those parameters doing those parameters are allowing you to optimize a penalty surface in ways that are very different than you'd see in more traditional machine learning applications you just call it say a convex black box convex salt or as a black box this is actually more similar to certain types of computations that you see in scientific computing Monte Carlo am lucky than MS computations and they're very non-trivial communication computation trade-offs often times so someone sounds like we'll need the full off book from you about the blurring sorry how did you have me back in the year now we're actually working on that it's not as mature as what I'll talk about today but have me back in the year and maybe it'll be mature enough absolute lend dips when you're ready to talk about it will tell you at the very top you know so spark back and and hear about it sounds good I think you're are you looking forward to talk all right thank you