Bay Area AI: Michael Feng, Product-Managing AI
Recording: Bay Area AI: Michael Feng, Product-Managing AI
so before I start I try to take a quick poll of the audience to gauge what level of technical or you know whatever discussion I want to have so maybe show of hands who in this room is a product manager or an aspiring part of manager okay developer data scientist okay and and who finally most importantly who here has seen the movie office space okay great great so I won't have to explain too much so so my motivation coming in doing this talk was really you know I've been building AI driven products last five years and if there's one thing I've learned they're really hard and now every almost every product now has you know is using AI or machine learning or AR automated reality so I want to make sure that I think the same kind of evolution in a product management community I feel like hasn't taken place so hopefully this talk will gives both the product managers and also those of you who are fulfilling that function within your respective teams think about it a little bit different so first of all my background in 2013 my co-founder Max and I started a company called dock lunch at the time we built a algorithm that would extract tables from PDF files so you upload a PDF file that would scan through it identify any tabular data structures and extract them out while retaining the structural characteristics of the table so the columns rows and so forth we sold ads banks along the way discovered that selling the banks is really hard like literally from proof of concept of getting paid with a year and a half so we change the product at something that analyzed documents and specifically showed them as web pages which allowed us to use JavaScript to track how long people were reading each page what things they clicked on and allowed us to deliver an application from to marketers that gave them more insight into their ebooks and why papers that led to us getting acquired by company called nitro where I worked with like we're Alexi with two scientists and nitro what answer we want to build a future of smart documents the theory being that documents are the inanimate objects that we use day-to-day but they don't really make a super productive so by doing things like detecting where the form fields are by changing the format so that they we fit a mobile phone or by extracting names of people places and gates we could make them more powerful so that's pretty good right haven't had some experience doing this this is the real top of my talk in five years I build three products of three you know for the three companies the table extractor it was like 80% accurate so it was unreliable such that you couldn't just upload a document and this habit automatically set tables people were really using at the first pass before working with their information in Excel the document analytics tool we had hundreds of thousands of reed sessions so people who read documents but you know that dad was really noisy and we clearly make sense of it and we couldn't deliver to our customers what they really wanted which was to figure out which parts of the content really worked and which and what could keep their customers to read more of their documents and engage more and finally in nitro we built an awesome form field detection prototype and a really cool thing that would convert your PDF and so you can we slow it for mobile device even if there's like two columns it was like a problem even though it's a fixed layout but ultimately by the time we left nothing had gone prototypes a product top productionize to a point where a paid customer was using it so you know in my book if you're not you know succeeding your family and so I still consider them failures even though I learned a lot from them so hopefully I can take in some of those lessons and and help you guys avoid and then take I made so today three things I wanted to talk about first is that why someone needs to wear PM hat why the role of a product manager in a in a data science team or someone a team of building AI products is even more essential than ever while the traditional product management process isn't a great fit for AI and finally for habits that I think that effective AI managers should have I know I chef Sevan but you know maybe a leave the next version will house oven so someone needs to wear the p.m. hat the reason I think this is true is because if you look at kind of like where the data sign commutes today it bears a lot of characteristics of an academic mentality and the academic mentality is one will give explicit problem statements both structured data and well-defined to sex criteria so if you look at the first one here this competition on taggle which is the basically predicting the question the question you are trying to answer is Cupid I provide a better way to predict and estimate home prices there's a lot there's a structured data set that they provide you which are all the homes have been sold all their features like number of bedrooms and four bathrooms so forth where they're located and the actual price that they were sold for and the security credits the success criteria is easy to figure out it's the difference between the predicted value and the actual value and that makes your job as a as a research engineer or a data scientist you know pretty straightforward the goal is also very clear is to publish a research paper that improves upon the traditional benchmarks the gold standard for this is the immigrant paper in 2012 which showed that compliment neural networks could you know vastly exceed the traditional benchmarks in the image classification tasks most of the papers are you know they're improvements of much more incremental but that's still kind of like the model in academia now in the real world a little bit different it's more like this you know your manager comes to you and he's like I want you to explore the data you know exploring and then you know afterwards you know let's get some insights or customers and he walks away literally I'm not joking this is exactly what happens in the real world and you're like what does explore does explore mean you know what in size we talk about here and the problem is that a lot of people aren't technical they think that AI is like this fairy dust just sprinkle on top of gap you know and keep sprinkling these insights will shoot forth like magical rainbows yeah and unicorns as well and that is not how it works right and so it's a product manager you're responsible for owning the following questions what problem are we solving why is it valuable do we have relevant data that's actually you know can might possibly answer the question how we're going to measure success and most importantly how do we actually build a system that works and this is why I think product managers are incredibly important in Nai this is probably one the most famous quotes from from office space but there's a grain of truth in it you know because essentially a product manager connects people who understand the customers salespeople executives clusters themselves with the people who are going to build something that solves a problem the data scientists engineers unfortunately right now AI this gulf between the customer people and the technical people is extremely wide and that's why as product managers we have to understand both sides in order to bridge that gap effectively finally most teams don't have a product manager right now so if there's no PM the leader team you're the one wearing the PN hat and it's your job to answer these questions so let me move on and kind of explain why I think the traditional PM process is not a good fit for how AI based products can be built so the traditional process is really more of a problem solution kind of like split and so it maybe waterfall I may be agile Kanban or scrum but but the unit economics the product management or all pretty much the same you start with talking users and gathering their stories find out what the pain points are then you define the requirements you know for what you want to build and how to measure that then you work with UI UX designers to develop mock-ups for what your feature or product might look like and after it's built you perform a substance testing to see in the national mock-ups and it actually met the requirements you set forth so to kind of illustrate why I think this is where this works and where this doesn't let's let's use an example of a product that can I they will use throughout the rest of this presentation and the the product I chose was is a product delve it's basically Netflix for documents so just like Netflix personalized personalized recommendations of movies for you based on your preferences delve basically the same thing but with office documents you know not as exciting obviously but so useful you know you think about you know a large organization where lots of people have lots of documents well be great to know the bunch of colleagues have shared publicly if it's relevant to you so a traditional product management process would be solving a problem like this hey users are not using a product right a very common problem pretty much every product has you know some form of this problem and maybe you talk to users and you realize well it's because you know they don't know what the products all about so let's build a tour you know so the first thing you see when you are using this product is a tour that shows you you know what you can do with the product a traditional PM process this is a good fit for a traditional camp process because you can say the story is as a new user I want to know what the top three features are so I can start using the product and the requirements are maybe are you can show a carousel with the features the user can exit at any time so you don't pick them off and it works well on mobile so you work with designers you get the mock-up and afterwards you know you assess did the requirements pass if so then you're basically done so any we've seen office space this has seen after they smash the printers in the field and they're you know they're very happy because they destroy some relic of old technology no but here's the problem let's say you still have the same problem head before users still aren't using the product so you go talk to users and say you know so what's going on and they say well honestly these recommendations are pretty bad like they're just not relevant to what I you know what what I want to see I don't like these documents or you have completely no relationship to what I care about and so maybe now you want to improve the relevance of the recommended content and this is where I think product managers struggle with it because you know we're very resourceful we can you know come you know we want to we solve things but the traditional kind of like problem defined problem create solutions have a approach isn't a good fit and like let me illustrate that by actually using a traditional p.m. process to try to solve this problem so maybe talk to one user the user says well you know as the user I want to see content produce like people that I care about right so the requirement is well okay let's overweigh the documents shared by the people who are in the same office as a user so you add a rule to algorithm-based expressing that you talk another user and that person says well I want to see content it's related to my job so you may you know add another rule to your algorithm that over ways documents shared by people in the same functional department and finally a third Jesus says well I want to see information around social events because I'm a party animal and so maybe you had a yet another role that overweights documents that are shared by people on the social planning can so I hope you see where I'm going with this because what these incremental marginal rules lead to is a bunch of heuristics okay and basically you know if you're in this area if you have this problem if like your model basically is a nested series of if-then statements where you're socially weighting these different factors back and forth and your model is like this gigantic block of spaghetti code and the problem with that is it's not a system that learns by adding more data you're not you don't have a up like a way to actually you don't know you don't know that more data is an improved accuracy moreover the more you tweak these algorithms listen you change one weight here you know that it becomes a zero-sum game because you are essentially you know pushing the model in one direction at the expense of another one and so if it's not a system that learns it's not AI full stop and this is usually I think a lot of people start with this because they really like all they have but I think if you actually understood kind of like you know some basis of machine learning in AI you can do a much better job and so let me tell you about kind of like what I think is a better framework for dealing with this all right I had like four basics four things that the iPod managers should be doing number one you need to understand the model there's a lot of content now about machine learning on Coursera Udacity this course off-task AI which I loved and frankly it's not very difficult for someone who has a basic you know university level math background to understand machine learning and if you don't do that it's really hard to make progress as a product manager building AI product so let me walk through example using kind of like that Netflix with document example can't see another show of hands who here is familiar with cloud result ring okay great it's about you know let's third to half the room so oh oh I'm not gonna send a lot of time on this but if you really want to know how it works just Google like a lot of info out there essentially cloud resolution is a is a is a recommendation algorithm that takes ratings of you know things that people submit anything uses that as the training data to predict ratings which have not been seen yet so and the key to doing this is that you construct essentially a vectors of latent factors so think of like these hidden preferences that may drive a lot of decisions we make but we don't know exactly what those factors are and the model essentially tried to tries to predict based on these weighting factors and then you know and train based on the error between the training data and the predicted readings so case in point let's let's take a kind of ratings let's say that there's three types of ratings you can give there's two which is a it's very strong interests one which is lukewarm and zero which means no interest and if the cells blanket means that is that the the person hasn't rated at all so it's an example let's say based on this this matrix what do I think the the rating that users three would give to documents the rate would be let me guess a widget Hannity think it's going to be a zero which means no interest one okay - okay yeah so it's probably as one or two somewhere near there but the point is it more like is we can actually do a systematic job of particulars and way we do that is essentially let's say we have four latent factors these factors may be things like in a proximity or released by the social community but more likely they under am album of all these different components usually there's no one characteristic but by multiplying take the dot product of the user vector what the user vector in the document vector we can really predict what that right here and by doing so we essentially have two matrices we have the training matrix and we have the prediction matrix and when we have these two sets of data we can essentially compute the error so essentially that's the sum of squared loss minimize that error by changing the latent factors and therefore that gives us a model to basically make this better so I'm glossing with a lot of details I'm hopeful that you may be either interested you can Google Cloud resulting after this and learn about it but the one key part here is that the more training data you have the better your predictions will be because you'll be able to check your predictions versus the training gap and finally we make the prediction you know I predict is probably very simple algorithm like show the user the top rated documents that he or she does not own and based on that then you can do more things so the second thing that I think all a I party managers going to do is number one is get labeled data because label data is what really drives you know these models and so some of you is probably heard the cliche data is new oil so I think this is actually wrong I think the really should say label data is a new oil now and also unlabeled data is the new dirt so label data are basically data which has explicit labels of exactly what you're trying to predict example you know things that are not label data server logs style images and files and s3 a storage system what is label data would be server logs where you also indicate which ones came from a malicious IP if you're predicting oh you know where the hackers are coming from another one is style images with markers of pools if you're creating a pool detector by which by the way is may think why would you want to do that it shouldn't companies find that really valuable because pools are a big indicator of risk and finally thousand three if they were marked which ones are confidential and which ones are not confidential that allows you to create a confidentiality clause supplier you know getting label data is probably one of the biggest challenges that a product or data science team needs to solve so a few ways to do that the the most common way is something called transfer learning where you take pre train models from another domain and then you may add a little bit label data at the very end to fine-tune it in order to improve on an existing algorithm this works really well in computer vision so it's I mean you've seen things like the there's a cucumber farmer who basically built a computer vision system that that detected cucumbers and non cucumbers he basically just used a pre trained imagenet model to do that the the Silicon Valley episode recently about come by where someone built a hot dog or not hot dog classifier I'm pretty sure you can just take pure image net and get to like at 80% accuracy just at one image net alone without any additional data another one is elsewhere slave Lee cloud flower and Mechanical Turk are the two most popular this is where you hire people to label your data for you finally well in addition I think brute force is always an option anecdotally I've heard that a lot of what Google has put out in terms of ml was built on the back of a lot of individual Google employees who spent a lot of time labeling data you know and the rule of thumb I always use is how long would it take you to get to 10,000 samples in your training set more often than not it's actually less than you would think because you know it's like 30 seconds per training example and you spend a few days on it you're going to get there and that sometimes that is the best way to get the data if it's you know if you don't have it already but sometimes none of these three really works and in that's when user feedback becomes your only option this is most most common when you're working with sensitive enterprise data where privacy issues prevent you from outsourcing it or boot forcing it or where it's something that is personalized where you want to personalize the output of something particular for a particular person and therefore you know training kind of you know training if you only have one person it's trained on it doesn't extrapolate onto other other people and this is why I think so the third thing is really important in terms of AI product manager which is yeah you had to construct a feedback loop so the feedback loop works like this you display predictions you collect user feedback you add that feedback back into your training set and then you refit the model I think the one thing that I would just be uh if you're going to do this is make sure your feedback is structured in the same dimensions as your training data and it's not structured that way transform and such as it's such it is so in the example of the Netflix or document application maybe one way that you can turn that feedback into all your training data is for every document shown you want to assign a rating of 1 2 or 0 so maybe one is if the person clicks on the document and opens in once if it's 2 if the user is opening a document for second or third time and 0 is if the document has been shown to the user more than three times and the person doesn't click on at all and so as an example in this case let's say the person click on the document in the top left the first time they click on it you look at your training matrix you plug a 1 into the blanks that there was that the blank area that was there before the person clicks on again you turn that one in or two and and I think you want to have that feedback loop in place on day one and the reason you want to do that is because you're essentially collecting label data label data is what fuels you know AI based algorithms and we don't have that feedback loop in day one you're basically wasting right data finally I think it's important to measure performance in relatives and not absolute terms that's because the first version that you ship will probably suck a lots of example to this out there this is Twitter feed I found called bad Netflix recommendation so it's like you know if you want you like the movie Something Borrowed you're like taken you know I actually think this is a I'm pretty sure this is Joe Sekou because you know if they are using club for filtering and they're using it based on rating factors the titles of the documents shouldn't have anything to do with it but I still think it's funny another one theory Theory you know windows launch was pretty bad you can argue it's still pretty bad you know and finally who can you know I think walk familiar with this pay chatbot that Microsoft produce which you know it was strange something that I think they did not envision on very first day in and so I think the reason why you want to touch well it's performance because you want to figure out how much data you need to get it good enough and most importantly you want to set expectations so when you're a manager basically you know or your VP or whatever status hey how's it looking and they don't measure how good the product or the algorithm is when you launch it because it's probably nothing very good but you can say yes it's not very good now but here's how much data we can we need to get to that good enough state then it's easier to make and a business decision as to whether to continue the product or to kill it so illustrating that in graphical terms you know with more data the accuracy of your model will improve the slope is your relative performance if the line that says good enough then essentially you know the intersect is how much data you need to get there so you know just putting out together you know here's my four habit first you have to understand the model because that's really the core of AI the second is you need to have a strategy for collecting labelled data third is a has a feedback loop on day one that points back into your labeled data and number four measure performance and relative terms and set expectations based on that that's it thank you okay um yep one way I help well I think you need a lot of label data you know because I I think that I think there's my sphere so far that AI works well in where your data your data samples are homogenous right so when you have we went when data point a looks pretty similar data point B to data point C and you have a good amount of them and so the the I think we talked about we think about like kind of like product creation there is every product you know every the process of creating every products pretty different from one another and so I think in order to the first step before you can start applying clinical AI type process is a way of standardizing across these different product formulation processes so you can apply some type of yep yep yep yeah so there's a lot of research into that I didn't want to get into exactly kind of like how cloud where collabro filtering doesn't work and that's exactly when the areas that doesn't work another one is just boost starting them from scratch because it cannot be a cold start problem so to be honest I think both delve as well as Netflix have moved way beyond Calabro filtering as a way to provide the best recommendations to people but and I think I think what I was trying to do here was more illustrate kind of the need for a non-technical person to understand the basics of the model in order to kind of like make progress but I agree with you that you know if if you even if the models really complex I think you need to understand kind of where where it falls down and where it works well in order to make the effective decisions you need to make oh yeah this one depends on what you're trying to predict so usually they're usually I think it's easier to get to know that upfront because usually this it's framed as a percentage accuracy level so for example when we were doing cable detection we we we estimated we need to be something around 97% accurate based on like the types of tables were pulling out and what people were doing with them so I think with any type of exercise like this usually there's some type of benchmarks you think you want to hit where people will truly becomes useful to them and where it becomes automatic enough there's also this also interesting enough depends on if you're doing a conformation model or a recommendation model so a confirmation model is one where the user assumes that it's you know it's it's good and they don't need to take action they only take actually correct it the other model is a recommendation model where you know you present three choices to the user and the user picks one of them so if you know the the confirmation model is let's say one where you know you're checking or stable detection I think is a more confirmation model if you're if you're uploading a PDF with like with 30 cables in it you don't want to go through and check every single one afterwards whereas a recommendation model is something like a smart reply in Google where it recommends a reply to a particular email and so the user can pick one with three or none at all and so like whether your confirmation recommendation of sex what Goodenough means to your algorithm okay yeah I'm not used to that actually cuz I think I think that one of the things that you know one of the reasons people do say you know data a new oil is there is - there's two schools and all they're supervised and unsupervised learning there's definitely use case for unsupervised in in business terms I've sounded more useful to sort of like indicate where to point your word word of supervised so as an initial first cut when going exploring data to see you know which which features you want to use or what types of labels you want to collect it also helps a lot in visualization so like the cluster charts that the mark showed earlier and to kind of communicate like what data is looking at large but the reason that a supervised learning is pretty much is most of what's been released today in terms of you know real business products it primarily it's because usually in order to produce business value you're trying to predict or classify something whereas in supervised you know if you don't have kind of like a direction that you can point the algorithm toward it's hard to iterate towards something usually the the first the first cut isn't going to really be good enough and you're going to work on this model and the data set for a good period of time before get them valuable so supervised essentially gives you a goal post to do that this one yep right yeah honestly to me in my opinion uh zero I think that I think that the frankly I think that we're still looking for the right terminology here you know the there's a school of thought that AI is sort of like in caps place NL machine learning and robotics and augmented reality and so forth deep learning as well and I think that they're all kind of like somewhat related but when I think of AI I really think of algorithms that can make decisions without being explicitly programmed so no defense statements like nothing that says if Ada and Ruby is something that can basically make decisions and the more data you give it the better decisions makes yes sure as long yeah yeah yeah so I think for I think for for these types of applications where you want to measure kind of like time spent and productivity is a you know you don't want to measure you want to say if I spend more time on this they're not but but I think that time spent does indicate an interest because most of the things we find out there like we only see a small subset of you know Woking what we can choose from so I think one of the things that that's why I use x clicked as kind of like the what I think it's benchmark should be because you know a large organization produces millions and millions of documents and you know if I you know indicate preference from one over the multitude out there that's available that some indicate indication of interest but I think the amount of time I spend reading that document probably is in a great proxy for how good or how interesting that subject is no problem exam on this one right here this one okay cool yeah yeah great thank you very much [Applause] [Music]