Devreal

Build Smarter, Don't Just Build

Event: Data by the Bay

data.bythebay.io: Hernan Asorey and Robin Glinton, Build Smarter, Don't Just Build

Recording: data.bythebay.io: Hernan Asorey and Robin Glinton, Build Smarter, Don't Just Build

so hi everybody my name is heran I'm going to be uh talking close to that because there's only one mic uh so it gets there uh the I think that the main objective is it's never a good place to be between you and lunch but we will have 20 minutes pack of 1 minute per slide approximately so we will try to go as deep as possible on whatever we can cover in 20 minutes so the idea is that um I was hired in the company about two years ago you got like another couple of speakers from cforce cell force is right now trying to invest a lot in intelligence in IQ in trying to use the data in not such an inspiring way like U it was presented before on health and things that were going to change people's lives but we want to change the way businesses have been conducted nowadays and we want to change the way we are interacting with those businesses and how those businesses are interacting with their customers and the way we're going to do all of that is just to inject in every interaction that we happen to enjoy in your favorite retailer your favorite Bank your favorite uh Communications company your favorite hospital is going to be powered by cell force and it's going to be powered by the intelligence that is going to be within C for so part of my role is thinking about the different personas that we have and this is just one of the personas the name of prod data science is exactly the idea of how do we inject data science in all the products that we build across cell Force cell force is a multitude of different clouds um we leave and breathe out of the log lines and what do we where are the log lines are the application log lines every interaction that you're doing with every product in cell for across all the different clouds are is being collected in real time we mind that interaction and we try to identify how the pro can get better and that goes from Hit maps on what the product is being used from what not being used that goes from uh features that are most popular within specific segments in the customer base um definition of the road map of a multi-billion dollar company that has to identify how to dedicate for the next three quarters 200 resources on development on the best ideas that are going to get the highest degree of adoption the way we do development the way we ideation the way we do even design thinking is now educated by evidence nobody just goes into a meeting saying I have a good idea it's like I have a good idea and the propensity of that idea to to hit a success is this much and the propensity of the customer base to adopt that idea is this much and this is how we're going to develop that was something that we were not doing two years ago so this is one of the personas and the Persona is a product manager so the main I mean as you can imagine we have tons of product managers across the different different clouds and it's not a it's not an easy task to tell somebody that is coming with a very strong mindset you know what this is a great idea but it may not likely be the best one for what we need to develop right now to unlock the next layer of adoption we all live and breathe for for creating products that are going to be highly used having a product that is beautiful and nobody used it doesn't stay right but the product managers usually feel that everything that they come up with is the best idea ever and they has to make it and they has to do it in the Sprint and it has to happen so we needed to create a way of changeing the dialogue in just making sure that that was going to be evidence driven that was going to yield a result that we are going to then translate into better user experience better user adoption which at the end of the day converts into you know the typical common dollar sign right but we all need to get to that step in order to get Del value so there are three questions in here the first one is like how do we help her identify adoption blockers if she's a product manager and she is trying to come up with something that is fundamentally new what are the things that and Robin is going to be talking about a project that we build that is called or DNA that we are borrowing from Life Sciences we're borrowing from everything that you can imagine uh most of the robing reports into me and we hire three people from galvaniz already um we have a very Mighty team of four people under him and that is the gist of this conversation because I told Robin we would used to work in the past and I let I let a team of 270 people and I said okay Robin it's like you work with me in the past with a in the tens of people now you're going to work with me with four people but you need to be as productive as what you were so we need to reinvent how we were operating and this is what we're going to learn today uh the the second topic and I just want to heit a little bit of different use cases is how do we help her identify and build the next best product or the next best in product this is all about Root Management imagine that now you are in a room with 200 PMS fighting to get that idea into the Sprint right and you need to say I need to create a Force ranking mechanism that is going to make sense that is going to be more of a organic graph that will have dependencies and the dependencies will be effort the dependencies will be likelihood of that feature to be adopted by a users the dependencies are going to be things uh like layers of monthly active users a metric that we leave and Breathe by is is how active the community is on monthly basis and you know well you know if I develop his idea versus her idea we would have more likelihood to unlock 10 million more M monthly active users in that particular app do you want that or do you want just more surface content so that is a dialog that didn't exist in the past and finally what is the right mix of product to build and here we're borrowing from basic old retailer right Ma and this is Market Basket is everything that we learn through many years and identifying what is the best basket of features that according to clasing and classification and other type of problems that we are solving we create digital personas that will equip the tool in recommending what are the features that have to surface through UI so you can use them most of the things are things that are going to be uh seeing live uh at reinforce so many of the things that um we're going to be sharing are of course able to be shared right now but you will see them live in product and this is how AI is going to take over how we present uh user experience so with that all right thank you Hernan um I I I think uh I I hope what you took away from uh hernan's uh intro is that our team has a lot of hard problems to solve we have to generate a lot of predictive models to help our product managers do their jobs um the problem is data science Talent is hard to find it's a cliche that the data scientist is the Unicorn the reality is uh that it really is difficult to find people with that really specialized mix of skills on top of that when you find those people they have a lot of competing offers from a lot of exciting companies um we think Salesforce is one of the best to work for but they have a lot of of great options so what do you do when you're in a resource constrained situation what you do is you try to find tools to increase your productivity and that's exactly what we did to try to make our small team more productive now what does productivity mean in the context of data science and building predictive models what it means is starting out we think of it sort of like a funnel you start out with a large number of data sets and sort of hazy ideas for models you can build and you try to funnel that down into the few sort of uh value driving or money-making ideas that you can actually deploy as models uh to do that you go through three phras uh phases you prep data uh you do experimentation um in this sense we don't only mean say statistical testing we mean any exploration of an idea space trying out different types of models trying out different parameters in those models and finally uh if you find a great idea it works you deploy something into production so in a reliable way uh uh you can get those scores or the outputs of those models right and when you deploy something you also need to to monitor it to understand if it's working um so what we did is through a combination of of procurement and uh in-house building we built tools to help us be more productive in each one of these phases uh in the first phase data prep uh we we found a tool that I'll do a double click on um that actually helps us uh to make our data sets easier to understand uh and find um and document in the next phase of experimentation we actually built one uh tool and procured another one we built one tool that helps us with uh large scale feature engineering to give our data scientists sort of a head start on understanding which uh data points are useful for the things we want to predict and we also uh procured a tool that we we think of as a continuous experimentation system that helps to try a lot of ideas in parallel and I'll double click later on what that is and finally in the deployment phase we found that we were spending a lot of time just trying to uh monitor lots of models um so we built a tool inhouse that helps us to manage that process by exception which is a Time tested way from from the days of industrialization to scale and effort so let let's first take a a closer look at the data prep phase um so in order to have a very efficient data uh prep phase of of building predictive models you need your data to be what we call durable it has to be able it has to be easy to find your data sets has to be able uh easy to understand the content of those data sets and the variables um uh from a business perspective you have to be able to uh easily recreate any data sets that you derive in that process if anyone's everever lost a team member who didn't document their work well uh you know it sucks to go uh recreate a data set takes a lot of time your data needs to be accurate um also uh a source of thrashing uh in a predictive modeling process if you find out that there's actually an error in your data uh also you need EV all the stakeholders to participate in documenting that data set uh you need the business folks it the data scientists themselves they all have to be a part of documenting that data set and finally you have to understand those sources of data so a lot of time is wasted whenever you have uh when your data sets are not durable if it's hard to find data you can spend weeks just setting up meetings uh to try to find data sets you can also spend weeks with business folks trying to understand what those uh weird camel case column names really mean from a business perspective um you can spend a couple of days trying to understand if the data is correct um best case if it's not you actually find out and you start you go back to the beginning and sort of thrash and lose time worst case you don't discover it until much later and it's actually causing problems uh from a business perspective still takes more time and finally you have to write the scripts so we found a tool called Elation that helps us shrink that uh week's worth of time spent on the data prep process into a matter of hours and what eltion is is it's a data cataloging tool but it does it in an automated way um so whereas traditionally it is tasked with uh documenting your data sources your data sphere this tool does it automatically um the great thing about that is generally it's a very timeconsuming process to do this what it means is it never gets done um in most data shops you just don't have well documented data you have some tribal knowledge at best a few folks that you constantly have to reference um to understand your data sets um so what happens is uh you plug your data sets into ation in our case uh we have a Hado cluster and an Oracle relational database and ation craws all those data sets and it cataloges uh the relationships between all the tables the column colums it does some NLP to actually give human readable names to the columns it also does rectification from a semantic perspective so it sees columns that it thinks are similar and it actually tells you that they're the same thing on top of that there's also a social layer where you can have write articles uh the folks who know the data best and users can write articles and have conversations around the data um so for example if someone finds a column that they think is incorrect they can start a conversation with the person who's responsible for populating that and resolve that issue and anyone who comes down the line can benefit from that because they can see the conversation so the netn net is uh the main benefits and speed UPS the data science that this tool gives faster on boarding you have an entire history of your data uh well documented data environment you can ensure your data is correct less thrashing and finally you have alignment on semantics so you can uh be sure you're using the correct data so at this point in time uh the the most important part that that we want to cover is that remember that remember that that slide that it has like days or weeks in two hours my my main problem is that now we are like close to 30 people right and all those people were being on boarded really fast and I literally had the experience of testing having a team with a or technology like this and without and I hired somebody that started on Monday and by Thursday with a tool like this she was doing her own analysis she was essentially doing data science in the more canonical way that we Define data Sciences which is creating the model but all the data syndication data annotations uh the way to search for repositories the way to identify variables the way to connect with people within the team that was at that point in time already 20 people it was all done through the social layer and is is definitely a tool that I will consider that we all know know that 80 to 70 depending on the on the field per of the time is spent on just data wrangling so this tool actually help us a lot okay in the uh roughly minute I have left I'll explain to you um this particular tool that we use for uh feature selection that we built in house um so we found that we spent a lot of time uh we're a product team so we have a lot of utilization metrics about which features folks are using a lot of data points um so we found we spent a lot of time doing feature engineering just looking through those metrics to find uh what was indicative of things we cared about so we built this system called orgdna uh our Salesforce deployments are called orgs so this helps us to find the drivers of of behavior within those orgs and the way it works is you can think of it as sort of an automated statistician it slices and dices the data sets looking for different subpopulations and it uses as a baseline uh the global po population with all with all the rows and for example a Target metric you might care about is revenue it takes as Baseline the revenue for the general population as it explores subpopulations it looks for uh it looks for subpopulations where the value of that Target metric is very different uh from the Baseline uh it does this you can think of it as F test um between the Baseline and the different subpopulations and when you find an interesting population uh the dimensions that uh describe that population are then your interesting features including interactions which is which is really powerful and this saves us a lot of time um doing feature engineering uh another Tool uh that we use inh house uh is a tool called Domino we think of it as a continuous experimentation system with an illusion uh to continuous integration in devops so whereas continuous integration in devops helps you to find bugs really quickly and to produce correct code this helps us to experiment very quickly uh so the way it works is it uh consists of two main components one is a very special uh Version Control System that's optimized for data science in that it allows you to version uh code as well as the data uh that the code is operating on as well as any social commentary and conversations around that data and the results of any experimentation all together in one artifact it also has uh a dockerized uh backend system that you can Fork different versions of this code and execute them um this Paradigm of doing dockerized data science is very powerful um and I'll do a double click and and explain exactly why so the reason versioning uh your data sets your code uh as well as your results all together is so powerful is it allows you to try lots of different things in parallel because you can make Apples to Apples comparisons uh this is the heart of any kind of science not just data science the ability to know that if you make a change and you get an improvement say in the Au of your model score that it's actually based in the change you made and not some you know not because you used a different version of a solver in uh in your logistic regression model for example so if you couple uh the Apples to Apples comparison with The Power of Docker um where not only is your do uh not only cuz you versioned everything are you assured that you're using the same uh data and uh uh that sort of stuff with the container you also assured that you're actually using the same operating system the same configs um the same R packages everything is exactly the same right so this means that you can uh use the power of Docker and the power of dist distributed system to try lots and lots of different variations massively in parallel and much more quickly uh find valuable solutions to your problems so one question we get a lot is why not just use GitHub well GitHub is not really meant for data science it means you can't version large data sets um it means that it doesn't have support for notebooks like Jupiter uh which are you know common tools that we use in data science so the the artifacts are really not around experiments and and that just doesn't work out well for data science so in summary the benefits we get from this tool are reproducibility uh the ability to massively try a lot of ideas in parallel and the the the ability to bring lots and lots of brains together with the social aspect to solve a problem so the last problem we faced uh was that we were spending a lot of time monitoring our deployments we're only four data scientists and as we deployed more and more models we found that we just spent too much time uh monitoring those deployments to make sure the models weren't drifting or the under uh underlying domain wasn't changing uh so we built a tool to help with that process so the tool allow uh takes in its parameters uh the type of model A forecasting model for example uh the type of metrics you want to measure that model against in this case it might be something like an rmse and also uh the Cadence with which you want updates about the performance of that model and what it out outputs for us is things like alerts um on model drift uh it tells us if underlying distributions of input variables are changing which is usually a sign that your do domain is shifted and it's time to retrain a model um and it gives us alerts of all these things so we can really manage our our pipelines by exception uh so the benefit of that is that with just a few lines of code um a data scientist can can monitor the models they deploy and spend a lot less time both writing the code to do this and a lot less time sort of scrutinizing individual models and with that thank you very much for your your attention we are at time the the deck will be published has a lot more information we're going to stay for a couple of hours for any questions that you may may have and I don't know if you have time for is this open sourced uh this is not open sourced this is not open source but but the vision is that we're going to be investing a lot on open source uh particularly starting sometime early next year but right now as you can see this is a an ensemble of processes and technology that is using a lot of Open Source packages and things that we we collaborate with but this is not open source thank you very much and Robin