Devreal

SF Scala: Stepan Pushkarev Interview

SF Scala: Stepan Pushkarev Interview

Recording: SF Scala: Stepan Pushkarev Interview

[Music] hello everybody my name is Alexey Krylov I'm the organizer of Bay Area a I met up and here we are on location at d2 IQ a very exciting meet up tonight it's about learning model management which is not the first thing which comes up anything about the eye and kill the robots and sell their cars but it's to be important and we have two Talk's today covering model versioning and model monitoring and so we have two talks and here also have step and push curve from hydrosphere and he is gonna be doing the talk about model more interest upon tell us a little bit about yourself welcome yeah thanks Alex see ya my name is Stefan I'm city of hydrosphere IO we're categorized as machine learning management company so we do everything from the model cataloging through modal versioning model monitoring and concept drift detection all that cool stuff that happens after the model has been built so yeah we're open source have decent coinbase so happy to talk and share our insights with a that such an intelligent crowd here so your regular at our meetups to come versus I think last you were skilled by the way last year and then skill both the buyers come in next week so what happened in that year yep good question good question so from the orange egg perspective so let's start with R&D yep we we were looking for a niche for model monitoring different use cases because like model is term is so such a wide term machine learning use case they have like hundreds and hundreds of machine learning use cases for some of the use cases models degrade very slowly for some of the models model degrade quickly there are texts give use cases image use cases we found ourself in more in text and image use cases so just because the cons drifter and the inference space is so big they're big there and the monitoring becomes a cool part of the model development actually mm-hmm we're customers if we can talk about like what are the typical folks you interact with yeah the the typical customers are the persona wise persona wise it's a machine learning engineers or machine learning kind of folks who are in between of info infrastructure quite in charge of providing the infrastructure for the actual data scientist or machine learning engineers or in charge of the machine learning the workload to be reliable and stable stable in production so there's this new buzzword ml oops this is the thing it's hard to define you know they I am not sure what's there Alda here it's more like just the thought leaders and some people who are thinking behind out-of-the-box and in inner terms that basically there is no such a role of you you don't you don't have a job a position at title of ml ops or data ops yet but I'm not sure at the end of the day everything is it's a software yes and like software engineering it's the world and yes it's the world so those are the same software engineers smart enough to recognize there is a need there is a gap there is a challenge and they just take it yes cool so I mean you really see so like I just came up with a kind of funny well short criteria yeah can you just shoot the machine right and so this is right and so do the people work it with you can they SSH into machine so yes of course yes I believe you know the the same way as you have people using Gmail slack and others yes like nowadays all the engineers use Python use Python is darker on a user level yeah as a of course they don't they kind of they don't go under the hood kubernetes to some extent a and of course you have they can SSH so well that's the question on Twitter it was kind of a joke and they're on the poll but some people get real life in it they say like this is like you know Tecna bashing like riot so I'm curious do you think if you're a data science in a start-up you should be able to so sage into machines do you see these people talking to you or you have engineers supporting that the scientists so I don't believe in drag-and-drop machine learning mm-hmm yeah you've seen the evidence from Microsoft Azure ml studio type of thing you've seen the evidence of Amazon machine learning I believe that service was it's also drag-and-drop machine learning of course there is there are there are teams there are organizations with the we can support who can adopt very quickly this drag-and-drop machine learning use cases but it's the the real the real you skills are they require heavy lifting they require understanding what's under the hood if you if even if you take this higher level services from AWS prints like it was forecast just timeseriesforecasting a black black and white bully a gray box of timeseriesforecasting it cannot use it if you don't know part of the timeseriesforecasting if you cannot use it if you don't know how to optimize your loss function and customize your loss function to to make this this black box more efficient to you so that's it's just a math and machine learning but it's just a case a real example when you when the drag and drop doesn't work mm-hm of course you have to simplify the things you have to kind of create create abstractions nice UI et cetera et cetera but it's not gonna be completely hidden from the engineer yes even like you know this kubernetes echo system it's so hard to maintain kubernetes it's like especially in the early stages and even right now like you you hit the wall every time but people and you have to keep that balance between kind of easiness and develop the programmer challenging geo you need you need to leave something for engineer to work on yes yes that's what that's the motivation of engineers here if you if you tell hey everything is solved you just need to preclude button engineers are like it's not a it's like a job security but they're advocates for better and more transparent solutions yeah I don't think we're at the level of iPhones muttering it's not an iPhone hours I like it if you need to get under the hood sound curious so you are with the startup which basically is a very critical part of deployment of AI because you monitor models yeah right and so interesting thing are reveals and stuff goes bad right so can you give me like an example so let's say the more like like data science is basically a mess up and redeploy a bad model what do you do what happens so the the workflow is pretty typical for any production workload TV if something goes goes wrong mmm-hmm you get an alert you have this like the whole pager ecosystem set up however the the crucial thing how would you generate this alert because the machine learning fails very silently no if it fails nobody knows what the result of money yeah the silent thing is showing your bank account yeah of course of course so the the most crucial thing is to sit somewhere in between the your bank account and the actual inference and generate your kind of insights alerts all sort of metrics I guess the yeah there there is like hundreds of hundreds of metrics that can be generated from a single of a super simplistic machine machine learning model mm-hm so and the workflow so if it's like expected behavior of the of your environment you can retrain it you can you can automate it in ideal world however what we've seen like the most the most kind of often use cases are just a training server training and production data mismatch no data skew bias data sets in training because it's it's not because like data science or machine learning engineers are did the bad job it's just because of the enterprise environment they the they took some data somewhere to do a data mining and machine learning model modeling and the world has been has changed since that time or that data data set data snapshot was not was not full enough to describe the entire entire world or did all they all they all the all the kind of inference space so that's and there are use cases when you just don't have physical ability to train on entire entire data sets so like the image image image data inference based for just just a use case okay so a construction monitoring or some like video video service or violence type of monitoring solutions very popular nowadays so a lot of computer every console every construction is different from from another like in different countries in different sites construction sites if you've trained their model on one construction site you deploy it on another construction site something goes from there yeah it is expected but you have some like ramp up period when you basically discover this edge edge in the German you cannot report people for instance yeah yeah I like there are constraints by country yeah yep so be immune like hardhat it's different from like electrical worker workers like the manufacturing workers others so they look different that but the the use case is the same and you basically try to adopt the same the same model for different use cases so this is a this is like a good example of like the expected expected drift drift Domon drift or just a concept dream mm-hmm there are unexpected when the big enterprises data scientist just throw the model over the wall mm-hmm and somebody started using it in different regions might be in slightly different use cases it just is just not working and they don't have any insight what's going on there in production for instance like the healthcare healthcare or farm are the if the model is being trained to to read a clinical notes and extract extract some like the information from there it's being trained for one particular domain for one particular like disease and or or just the vocabulary and it's being deployed in two similar but slightly different different clinic or or use case that you may not even be aware of so you need a chance to kind of get an insight that something is easy it's different there yeah you have it take a look now this is really cool this is like a really fascinating and of course you're gonna you know do your main talk and tell about it so but you know like for this I'd like to kind of wrap up with a question so you see you know like it touches all the parts of the business it touches business metrics it touches Imagineering etosha data scientists so next year all right it's kind of sphere evolves who do you want to see more do you want to like as your customers do you want to see the machine learning inertia guy do you want to see actual data side and the depends how much tuning I'm gonna build like is that the sign is gonna be able to you know interact with you is the business owner who as soon is losing money making money he's even able to see a metrics like where do you think is going first three people so it's converging into more a business application rather than just a tool for machine learning because yeah of course it just save some saves money mmm-hmm the however the the users are right now our machine learning engineers or make ml ops yes yes whoever it might be yeah but the the kind of post-production analytics of your machine learning model performance is something that business business wants to see this is how your models across the different different versions or in different different methods has been doing over the last few months for instance how what was what was the kind of a product KPIs and business KPI is there so III can see it more as a business application rather than just a tool for machine learning awesome I'm super excited for further sphere and for you and really happy to have you speak Adamo top and we're gonna be looking forward to a for progress you