SF Scala: Evan R. Sparks Interview
Recording: SF Scala: Evan R. Sparks Interview
[Music] hello everybody I'm Alexa grab roof I'm the organizer of barrera a I'm atop and here we're at night row which is hosting us for very exciting the table with determined AI which is a new startup and Evan sparks the CEO is here with us he's gonna give us a talk about this later which you can also find on functional TV where everything goes phone call dot TV but first of all welcome Evan thank you thank you thank you very much yeah so determined AI we are building a workbench for machine learning engineers and data scientists who are trying to get their AI powered features and applications out to market quicker than they could before so if you've got tons of data lots of GPUs deep learning models that you want to encode it to your favorite image classification or or text of speech kind of applications and so on that's we help you get those things built faster than you could before so what's our engineer first and then a scientist later right so explain it to me when I see the result of folks who say that they're deploying models to production and there is this whole industry emerging why why do we see this phenomenon why there is basically a whole class of companies which say we'll help you perfect answer I was just not another kind of software that you know a lot of people know how to deal with yeah so I think that what we're seeing is that fundamentally the way we're developing software today is changing as a as a result of doing data driven application development things change you know we've got a lot of uncertainty we've got statistical randomness we've got noise things that you have to deal with that are quite different than writing down a bunch of rules and software and so we now have the opportunity to make better decisions on the basis of data build models out that that can make more accurate predictions and so on but it means that almost every stage of the model development process it and software Duvall process is going to change so everything from versioning now comes down to more than just software it also comes out to the data everything around see ICD and deployment and monitoring also changes because the world changes around us our models react and they're more difficult to debug and so on than traditional software so every stage of that pipeline has to be rethought thanks to thanks to these new methods that are powering better applications mm-hmm yes you know I think here I was pushing this a little bit too to the data scientist so so here's my kind of technology so the the result of new folks who came into the industry and they learn both the designs and programming in Python and they do not go through the traditional career arc of a software engineer we learn the ICD and test-driven development and all this kind of stuff right so you know in my mind that's my kind of first explanation to myself that this industry arises because there's whole bunch of folks who are either data sciences with also 3d nearing experience all that business folks so the model is serving business so business people are watching it right the data Sciences are working on it but they don't have necessarily rigorous experience of deploying and monitoring roll it back and so kind of models intersecting so engineer data science and business people is it accurate you might or yeah I think there's there's one way of looking at it where it it says let's take you know these data scientists who are running wild with there are scripts and there Python scripts and so on and don't know anything about versioning and let's let's instill best practices software engineering kind of techniques on them and tell help them understand what it means to put things into production and what uptime really is about and so on but I think it's also more nuanced than that I do think that you know the just the the process of developing an application is quite different than the process of developing a traditional application it's a traditional application development you know you hit the up-arrow you hit make on GCC and and your artifacts build and hopefully your tests are run too and you get all green lights and you're like ok I'm ready to go here you hit the up-arrow you hit go and you wait for 48 hours yes your model to converge all the way yeah and you don't if it's actually gonna work it's harder to think about what what success is and what failure is and so I think that in in part because of that you've got a lot more tooling that needs to be developed to make the process of developing these applications much more pleasant so a little bit you know if you know kind of what is your like what is your dissociating point right I know new companies in the space they come up regularly so what is your key strength and what is your core like you know core competence which makes you think you're gonna succeed and kind of winning this in the space yeah so honestly like the team is is a huge piece of this so we've got a bunch of people from PhDs from places like Berkeley Chicago NYU X schoolers people who've been at places like Pinterest and so on so really strong technical engineering team experts in distributed systems so my co-founder Neal Conway worked on mesosphere worked on on Apache Bezos I myself worked on Apache spark and I lived so on as well as my my co-founder me tal Walker Amin is a professor in the machine learning division at Carnegie Mellon so we've got a really strong core of people who both know the distributed systems side of things as well as these scalable machine learning methods so that's a big chunk of it from a technology perspective what we've done is is take a bunch of the research that we we've been doing over the last call at five five plus years around scalable methods for ml and and think hard about how do we commercialize these and how do we scale them up to the new new world of deep learning and so we have some proprietary algorithms around a type of revenue optimization I'll be telling you about sort of the open components of that later tonight we have some really sophisticated techniques around cost modeling for distributed training and deployment and so all of that kind of adds up to an end-to-end experience we think in that hopefully our customers think results in in dramatically faster times to build and ship applications unlike you know you've ever seen before nice nice I mean real excited that you know you met them that the citizen component because often in kind of traditional AI talks we we see a graph we seen on a notebook I don't see the disturbance distance behind this and I'm really excited that you know that that comes with the forum because in a real production setting you need any of the six release yeah it's it's really hard I mean it you know if you're talking about waiting you know days or weeks for models to converge you are starting to think about one of the natural things that you think about is how do I throw more GPUs at this than I then just the one at one at a time how do I do that in an efficient way communication efficient way how do I get scalability that I wouldn't get otherwise from from sort of the the office shelf tools how do I do that in a very principled way how do I make it fault-tolerant how do I make sure it's versioned and so on and and that's where sort of the system's expertise starts to come in and I think that's a really important component of the experience yes so this is great so I know that you know I'm blab actually makes the same kind of connection right but maybe for the reviewers who did not read up on that much you can tell us you know what makes us special why you know things like spark came out of it yeah so the Empire business terrific place at at a point in time where you know you started to see an industrially this trend where datasets were getting bigger and bigger it had never gotten it had never been nearly as cheap or easy to store and and process a lot of data but then the question was like alright we've got all this stuff sitting in HDFS what do we do with it and luckily it turns out that more data is great fuel for machine learning techniques and so the the people who who founded the amp lab my adviser Mike Franklin Mike Jordan who we worked with in the past and and a bunch of the other you know faculty there really saw this issue and said hey wouldn't this be a great opportunity for communities that haven't traditionally talked to each other you know distributed systems databases operating systems folks and really hard core deep theoretical ml folks put them in a room for five years and get them talking to each other and this is you know probably 70 or 80 PhD students you doesn't a dozen or more postdocs a dozen or so faculty members sat there and really started to work through these problems and you know I would say that it at the beginning of this this period which is when I kind of joined the ampulla things were really just getting off the ground and people were kind of uncomfortable talking to each other John du cheese sitting in one seat Matea Hari is sitting in another seat and both professors at Stanford now but it was like they never would have talked before and suddenly they got together and and more folks like like them got together and started thinking about how did the how do we reconcile these two worlds and and I think the net effect was something much greater than then either of the two sides would have come up with on their own and it stood for those machines and people yeah that's right people was absolutely absolutely a part of this and and both sort of the people who are thinking about these things and using the systems and that was important but also there was this this really great flavor of crowdsourcing and how do we think about how we leverage you know the ability to collect human powered data from lots of people at once as part of this and so there were some really interesting projects around mobile devices around using crowds for data labeling and so on that that started to come out and you know it was a great line of research so this is kind of humorous semi humorous observation about the structure of all employed startups okay and if you go to their you know if you go to the talks at some point they show you a chart of the architecture and I call the three layer and black pie it basically it has three layers at the bottom of all things the middle layer reads from the middle layer at the startup it may be spark and I be a Luxio or the mesosphere all right so it sees on top of stuff so tweets from are things like HDFS or s3 then the spark and then there is a whole bunch of apps like in my lab or in application layer frameworks which is on top of the pipe so if you look at every implant you know startup deck or slide presentation later by so what is your three layer part yeah absolutely so sometimes we call this lasagna diagram okay but yeah so our three layer pie and you're right there is one of these things and I was talking to a someone over lunch today and I said you know as computer scientists were flagrant we love the inventing layers of abstraction that everybody else should use and so we're guilty as charged but our familiar PI we start with the under layer of of course distributed storage and compute framework in this case mostly we're talking about GPUs but also you know things like edge devices that mobile phones and so on that might be deployed at the edge for inference on top of that sits both the existing application frameworks for deep learning so tensorflow cafe Karis PI torch you name it we tend to be we we try to be as framework it's not agnostic as possible there are dozens of these frameworks we can't do every one but we support all the popular ones that we can okay and then alongside that we have our stuff and and above it which is a distributed job scheduler something we could think of as the the trial optimizer a hyper parameter optimization set of techniques which we call the experiment manager and then a metadata story that really traps all of the things you've ever tried about you know developing your deep learning job over time and then the front end is really sort of command-line interface set of API is to integrate with the system and the a user interface to help users kind of understand how how are my models doing over time cool cool so it's sort of sending to know that you guys have supported multiple frameworks and so that that's so like do everything kind of pilots in the works can you tell us a little bit like how's it going yeah things are great so far so we company's still early stage we we were founded late last spring and raised some money over the summer we spent spend a bunch of the time kind of building on product and so on and we've got a few paying pilot customers people in healthcare kind of the genomic space as well as as well as IOT and talking to some folks in autonomous vehicles and and financial services as well so so far so good we've got a few real users kicking the system every day and giving us feedback and so it's been good to see very cool so in those missions what we're doing to be the year where as a little chicken was really a world a world domination obvious muscles you know I think that there's really this opportunity to help shape the landscape of how people develop deep learning applications and you have these great application from work frameworks the tensor flows the carats of the world that are really good it's sort of model training on single GPU but there is so much more to the application development lifecycle around data collection around data set management cluster resource management deployment engineering that needs to be figured out and there isn't good software there for it today and I hope we can contribute in some small part to that and and make the world a better place any wasteful community to chip and then open source kind of points so so far no nothing for open source from the company we've certainly been talking about which bits of this should be things that belong to the community we we haven't announced a formal strategy there yet certainly you know working on making the the application frameworks better is a is a big way that that people can contribute so I'll great well I wish you good luck in looking for as you talk great thank you very much appreciate your time thanks you [Music]