Scale By The Bay 2018: Fireside Chat with Richard Socher, Chief Scientist, Salesforce
Recording: Scale By The Bay 2018: Fireside Chat with Richard Socher, Chief Scientist, Salesforce
you all right all right guys we're about to start so if you can get out on the back take a seat this is the rather unusual for us we kind of do panels we rarely did first such as last time we did them during the happy hour and everybody was drinking so we didn't get as many people we actually employ it on the cordilla accordionist with a conga line to lead people back so now we decided a little bit earlier so thanks for being us here so this is Richard structure he's the chief scientist of Salesforce he is a big friend of by the bay conferences he started text by the bay in 2015 at meta mind and he was a data by the bay and I think I think in between basically was a story where multiple companies were doing data by the bay at least to the board by Salesforce so so it's a very good thing you know if you submit talks to to by-the-by events you get bought by Salesforce it's a it's a tradition now so encourage that and and so I think it's it's really fitting right what what we did this year and last year we brought a lot of topics from data to scale by the Besser this used to be soft rain you know conference but now with machine learning at the end of data pipelines and AI write basically machine learning is not called AI we have essentially you know with connection so Carlos gastrin said that you know every applications a machine learning application right so he was saying it for a while now it really happened so so Richard I think is a great example of kind of cutting edge industry going into academia going into industry and I think it's really interesting how this happens so we're sure maybe a tell us a little bit about the career you know what you were doing at Stanford that meta mind and how you went to Salesforce and very no all right well whole career all right so I guess I'm originally from Germany I came come to us for my PhD and then thought I was going to become a professor actually accepted a faculty job after I finished my PhD but then I thought well maybe for one year I will do this startup and then maybe I can transition and become a full-time professor and then of course in one year not much happens in the startup and really I had the team at that point and I really couldn't leave the team and I was also realizing that you can do some amount of research in AI in a startup though I think probably not for more than a year or so then you read me to mostly focus on products and customers and sales but I was also then starting to teach at Stanford and I loved that for three months of the year but I don't know if I would have wanted to do it nine months of the year it's a lot of work and not all of that scales in terms of student interactions and so long story short stayed at meta mind and then in 2016 we got acquired by Salesforce and it's a really good combination because in order to make successful AI products and actually you know be sustainable in AI you need to have three things you have data good algorithms and an understanding of the workflow of how you actually incorporate those AI algorithms how do you actually change how people do their work or do whatever else they do and so in Mehta mine my startup we had some of the best algorithms but really as a small start-up and the very academic team it was hard to get companies to give us the data and then help them label it and so on and now I guess we're a little bit ahead of our time just like others in terms of labeling the data at a lot of startups now that just label the data for you and then focus on maybe just self-driving cars and helping companies label across different modalities for South Korean cars are working on just medical image labeling and things like that and then the third thing you need to really at this point I think the time is over for building general purpose AI startups and you really need to be laser-focused on specific vertical really understand if you want to change the way service works you want to build a chatbot that automatically gives answers to service requests you need to be hyper focused on just that just chatbots just for you know financial industry banking or something like that and help people like you know recover their passwords for their banking accounts and find out how they change or open a new account or something like that and that was kind of what sales was really good at because we run sales service marketing and have a huge scale in that and of course also have the sales force to actually get that out the door so it's a perfect combination of the algorithms being put in the right place nice so what was the kind of best change going from academia to industry what was the worst change that's an interesting question so I think I think the scale actually as were right with the theme the scale that you can have in industry is just unprecedented as best case scenario as a professor you may have 10-15 PhD students and beyond that you're really just a bad advisor because you're not giving them the amount of advice that they really need so you can't really ever scale and then of course you can't get the stuff into production so when you think about how do I have the most impact in life with in my case artificial intelligence that I'm really excited about I really want to push the state of the art forward and do academic research but then I also want to see ideally they impact more quickly in real products and I think AI and I don't think that's too high P AI will change pretty much every single industry out there and it's just a matter of having right data set so if it's agriculture medicine CRM like banking like everything will get changed with AI and so it's just important to think about how to do it so basically what I'm most excited about right now is that AI is in this interesting time where you can actually do fundamental research in companies and we see that the bigger companies get if you're a small start-up you can't do it for too long right you can only think so far before your funding runs out or you need to think about scales as sales but once your company once companies get bigger and they start thinking about a three to five year horizon you know do we still exist and are we going to get disrupted by other players and when they start thinking longer-term then they start having research groups and then you can actually do fundamental research in in an industry and that's the case certainly for AI right now that's why it's so exciting to be there and then having like really great engineers to collaborate with and get those products out the door is really fulfilling so yeah I found you know you guys really regularly publish and you're on the cutting edge of deep learning research itself alright and so I'm really curious so it's not often you know that the research lab in the big company has the autonomy because you know there are companies like IBM who are used to run research centers our sales force is not known for doing that right so I'm curious how did you manage to keep this balance and and was it kind of the wheel from the top kind of to do it this way did you have to fight for it how do you manage to kind of do both things you know advance the frontier Phii it also protect eyes good eye I think it has to come from the very top of the company and I think you have to have buy-in from the CEO and then ideally also you know the CTO and in our case or two co-founders Parker Harris and Benioff and then Trini who runs all of technologies so once you have that but then it's also it's a new muscle and a lot of companies don't have that muscle yet and only some of the really large companies like Google Amazon and so on really have it and have it flexed enough to really know how to use it and so the way I tried to go about it is that I knew that if I just did pure research and just published a bunch of amazing nips and icy mail papers like I would probably not be able to grow to hundreds and hundreds of like pure researchers and so I tried to make sure from the get-go that we have some people who are really excited about just pure research but then we also have some people who are naturally personally excited about applied research and actually AI engineering in the group and as long as both of them do their jobs you can grow up together and then we actually have a couple of different ways we interact with the company and to be honest I don't think we have it fully figured out for all eternity but there are sort of four different ways that I think makes sense for research group to interact with the rest of a company and we're kind of doing all four so the first one is we just give advice people know already they know how to run a model how to actually have something in production but it's you know 65 or 70 percent accurate and they really want to be 90 percent accurate and I've had some amazing collaborations with sales teams at Lexus there and others where would that worked out really well because they were already quite advanced but then would also have some folks who couldn't really take the advice and implement it themselves because they don't know all the papers and don't have the background and so for them we actually sometimes also send them code we train a model and especially for sort of prototypes and maybe testing and trigger durations in the beginning you can just have the model train already and just say here's the prediction or testing code so not testing in the sense of you know unit tests the testing in terms of machine learning testing um so predictions and then the third thing is we tried to build api's because once you do the first two a couple of times you realize okay well we've given this advice to that group and now we've given the same advice to the other group and that doesn't scale because we have like 33,000 emplyees and thousands of engineers and so and likewise if you just send them code you know some people might ask for similar code you know I can now they're building and running infrastructure for that code then the other groups building the same kind of infrastructure and similarly so then you start to see the synergies after doing this for a while and then you want to create api's and I think the API model is quite powerful because we see a lot of synergies between different groups like we want to do for instance sequence tagging trying to extract dates and times and names and company names from text and it turns out that super you know for chat BOTS that's super useful and sales emails and that's super useful on social media and in a lot of other places and so once you have one API that's really powerful and general then it can be used across a whole host of different applications and then the fourth one is actually when we create a completely new kind of product which it turns out if you have a large company it's actually kind of fun because there's very little friction with anybody else because it's a new product line and nobody says oh but we were working already on this product why are you doing this so one of those new kinds of product lines that came out of research residential voice assistant where you can just walk out of a sales meeting or any kind of meeting in the future and dictate what you know happened here like oh I just had a meeting of Acme corporation and we should create a follow up with their VP of engineering in two weeks and maybe upgrade this opportunity to be like half a million dollars and then you just say that and then we do some voice recognition natural language processing sequence tagging and things like that and then we also have a deep connection to the database of all their customer data and we can update that entire field of this um and that's a super powerful new application basically the least fun part of using most enterprise software's data entry and you make that data entry super easy for people so those are kind of the four different ways code API is advice and alright so I want to look really nice Alexis Rose who gave the talk this morning with his colleague so here you know if you seen this song it was great right so and so I just want to pick up right on so that's you know a problem which you know we can all understand and if you have the contact graph in the continent 'ti and you want to recognize is this the same person a lot and and I'm sure we have a protocol email insights right which which it's using can you talk a little bit about how that kind of collaboration goes sure yeah so Alex this was actually a great example of where we gave advice because they already knew a lot of things that they wanted to do and they knew how to do them but maybe you know they're all use it of course if you're using doing natural language processing at this point you would be using neural networks of a sort but then neural networks have a lot of different knobs to tune and you can get a model that works reasonably well like an Ellis TM but then when you have a research team that's really looked at LCMS in like all the gory details and all the various options whistles and bows and Russell said you can tune on them you realize like maybe you need by directionality you can have various kinds of dropouts and I don't want to go too technical but you know there's a ton of different ways you can basically make it harder for the model to do well and during training you can delete certain weights you can delete certain activations as the neural network is going over text you can delete certain words you can have certain deletion masks that you keep replicating they're a ton of variants like at least 10 different ways you can do drop out on recurrent neural network sequence models and so and then they're wholly completely new kinds of sequence models I quasi recurrent neural networks that are much faster paralyze able and GPUs and things like that that we've developed in our research group and we basically were able to work with them and evaluate a bunch of different kinds of sequence models and help them improve the accuracy by like 20 percent so there was a super-fun collaboration where everybody sort of came out and felt more empowered to use the money so this also you know it's like an engineer's dream right because I remember you know I kind of fall more on an engineer site when it gets to systems so my favorite algorithms basically is take a huge array sort it and take top ten right usually works they just need to know what to sort on and so right there so let's say your kind of classic engineer and you are tasked with you know doing something like an old P right so kind of you know your normal mode of operation is supposed to scientists right like you don't look at the top researcher you see what open source is available how quickly I can steal it is it you know building right can you plug it in is it fast so so what normal help is like it was extreme or like some existing like Engram package so you use the basic common sense right and usually it will work and so what forever remains unanswered is could it'll work better right and so so in your case I think it's a unique resource that you know engineers can come to you and you basically know what's going on an LP and you know the state of the art I'm I'm wondering right so do like several like more folks like a lectures come to you do they already have like the home cooked and LP setup and are you were able to achieve speed ups what students right in this kind of because normally in companies do it bottom-up right they take stuff from open source they don't have you know like a global Authority amenity in their ranks so so what do you find in this interaction is there a pattern do the general do okay or should we just all like you know turn to experts like I will all underperform and dramatically on our LP stop it is true that anybody actually I would love to hear questions from the audience too so if you have a question after this one just you can raise your hand maybe we can open up also make it make it exciting and answer you kind of pressured about AI and all the things attics applications and so on but I think it's easy at this point to call yourself an AI startup or being doing no doing AI because yeah you just download some tensorflow code and run it and do some data pre-processing and boom you're like an AI startup and so the and in some cases to be honest if your application is just very straightforward and you're not trying to do anything novel and you know that's not redone before like self-driving cars or something like that you can probably get quite far with the standard tools like in and I love the fact that AI has become more and more accessible and you can learn I actually have a lot of folks who learn through the the Stanford class that I've been that I've taught the last four years and all the YouTube videos are online you can basically go and like cs2 24 and not Stanford you and learn and that Stanford class really goes and the videos go from like what's a word vector how to represent a word and learning all the way to like the latest research results and so it's actually quite incredible how much you can learn these days but then you know it is still a huge effort and of course you need to know linear algebra and probability and statistics and Python and PI torch hopefully and probably tensorflow still in a lot of places so it's not like it's super easy but it is actually if with the right amount of will and background and time investment it's quite accessible these days so we do have a large number of customers though who don't have the people who even know to go there and have the drive and will to do it right we have like non high-tech companies also that are our customers who don't have large multi-billion dollar R&D budgets and things like that and so we we do give advice also to to our developers and to our customers and we usually when it comes to AI which i think is actually a reasonable strategy we first start with packaged apps where we just say all right if you just want to you know understand your email it's better we'll just use our standard email client and we'll extract the emails for you and suggest follow-ups and suggest calendar items and things like that or you know on social media you want to know how to do social media listening just like it comes with sentiment analysis like marketing cloud social studio or I'm gonna find your company logo on all the things all you know tweeted images then it's just like pre-trained on the majority like of the top you know 5,000 or so company logos so there are a bunch of things that just work out of the box but then because we're a platform company it actually becomes harder for us to do the next couple of steps which is pretty much an enterprise software when you're not the customer because you're not giving up your data but you pay a lot of money for that software you have a lot of requirements and the requirements usually are I want to change I want to change what how this works what kinds of inputs are given to this algorithm and then people have custom fields and custom objects and custom workflows for their organizations and special you know it's company internal lingo and all of that and so basically that leads us and to build a platform so they can do a lot of other things and in AI in particular that actually becomes quite tricky so one of the things we one of the products we've announced I think last year at Dreamforce or this year was Einstein builder where you can basically predict one column from any other set of columns and that sounds kind of boring but that's basically 80% of machine learning and AI for enterprise companies like that column can be like how will this person pay their loan will this person pay their credit back on time should we give this person and insurance yes/no like their ton of different things in all kinds of verticals well this should just student you know will the student finish college if you're nonprofit should we send an email to this donor if you're and some NGO like anything can be in this column right then you might want to start predicting it based on the set of all the other columns that you've kept track in yes um and so that feature it's super powerful but boy it's also extremely complex because maybe people have like 5 rows and then they want to predict something right and they're like oh it's not working I gave it these five examples and and so in some cases you have to say ok we just don't let you even train it you need to have this minimum number of examples and rows to Train but then they're like you know they select rows and select the feature that is essentially just a modification of the row you're trying to predict so for instance you might have a one feature that is a dollar amount and another feature that another column that is just like if dollar amount is larger than this put it into this bucket and then and now you're trying to predict buckets but you gave us input the original continuous value of the dollar amounts in that case you don't really to classifier you just basically just maps that one dollar amount to the bucket and you know our machine learning is powerful enough to learn that and so you have all these confounding factors and and so on so I could go on forever but there's a lot of complexity when you allow non experts to just with a click interface to train this and in the worst case it actually all works but in this kind of an unethical classifier like they edit a race or gender column and then say should this person get a loan and now because there weren't that many woman and you know the 70s you started on business to start a company and wanted a business loan they're like oh who dad will not be successful they will not pay back because we haven't seen an untrained data and then it only woman loans to start a business and so you have to have all kinds of ethical considerations and again because we don't have access to our customers data you can't just like say oh I look at it and I don't allow it we can't look at customer data so we have to train them also to have this ethical mindset and realize all the potential issues and you might say oh you know we don't have a race column so no problem for housing but then maybe you have you know an income class zip code and the United States has tons of segregated cities and then it still picks up the race from those two other columns so it's it's a lot of there's a lot of complexity when you have this platform company that tries to really lower the bar for machine learning that it's not even coders but it's like complete non experts and who know how to build and you know and administer an interface but don't really go deep into the code well this is the first time I heard about that and it's I think I find it fascinating because you know we hear about things like Auto ml and the algorithmic approaches but in your case you have all the customer data like you have a federated data right your Big Data in terms of aggregations a lot of medium-sized and small enlarge data right but but now you because you have all this data you can help them train better but then you Kraft educate them which is you know I think it's a fascinating problem it's it's a good problem to have it is and which is one of the reasons I'm super excited to have with as part of our AI strategy we actually have trail head which is a learning platform where you can go and for this particular feature we actually want to force before you can turn it on you have to take a trail head unethically ice so that you're not accidentally building unethical AI systems nice so if you guys have questions we'll have for a few minutes please you know maybe ask a question and as long as you can and I repeat it for everybody else or we can bring your mic that's a better technology [Music] oh yeah and the question is how do you detect machine so there there's there's no silver bullet to this we're like oh we just do this and then you're done so for instance like maybe it doesn't make sense to have a gender column for in a housing and loans but if you're trying to find a new drug that deals particularly a pregnant woman I kind of made sense to like not or you know you're trying to sell breast feeding equipment or something then it doesn't make sense to go sell it to men and so that's okay it's okay to have bias in that prediction system right and so you need to really educate people and then of course there are some tools to that you can show people like these variables are very highly correlated with these outcomes and I actually have to give Google some credit there they have a lot of phenomenal tools to help understand your algorithms a little bit more and depending on which algorithms you use you can also immediately look at the weights like if you have a linear classifier and it's you know classifying for certain categories you can look exactly at like this weight has a very high you know impact on being in this class and you can help people that way and then we know in some cases like if there's gender and race in the database that I think for the most part I'm thinking of race but gender for sure because there are certain products at your weekend wanted market only a semi certain people and so when we see high correlation with that column then we can also give warnings and then in general for things where we don't have because you can have free custom fields that we don't even know what they are and it can be in different languages and whatnot in those cases we have to rely on the ethical training to help people understand it so Kathy Baxter was on my team of a chief or not she but the AI architect in the company and my team she has this great quote which is ethics is mindset not a checklist and that's kind of where we're also trying to go so it's I love that question because you gave good examples of how it's not slowing down actually like things got so much better in NLP if you've tried a translation system like maybe five years ago it was just laughable like you could not use it now the best translators in the world almost always use a machine translation system to do a raw translation and then they just clean it up a little bit because you know there's some sort of common sense things that AI still can't have and in some ways MLP I would call natural language processing is sort of AI complete like you need to know everything about the world and the visual world at common sense and reasoning and thought and and all of that to really fully solve all of NLP but in terms of real applications like summarization we've made a huge amount of progress in summarization basically up until last year when we published one of our summarization papers they had not been a single paper that could generate multiple sentences and it would be kind of coherent even the best translation systems that you use in google translate they just generate and translate one sentence at a time and so that was kind of a big breakthrough now bird is another good example it's another way to pre train work vectors and basically use unsupervised text as your supervision signal by trying to predict the next word or the next sentence or words and in the context and that's a very simple idea that we've seen in a lot of other and the word vector models and so on but it's really powerful and it turns out captures a lot of common sense reasoning too and so long story short like we've seen a lot of improvements and yes an LP unlike vision vision for many years was in a state where everybody was publishing papers but nothing really worked right and i'm like guilty of that too i published vision papers back in the day and like graphical models and so on and it was just like the accuracy you would at the state of the art from 33 to 37 and i was like not production-ready right and then deep learning kind of an image meta thing had the first gem by like 20 percent or 15 or something right on the first version of image net but NLP is already a multi-billion dollar industry it was already kind of working reasonably well in search and an advertisement and made billions and billions of dollars which isn't still not the case for computer vision there's still fewer vision applications than an LP applications and so and again NLP is even harder it's you know in some ways and there are lots of funky arguments you can make but I would argue that language is kind of the most interesting manifestation of human intelligence like a lot of animals have reasonably good visual intelligence too but no other animal has such a complex language as we do and so long story short I don't think we've hit a barrier yet I do see a lot of improvements in NLP and deep learning and I also think that the notion of deep learning has changed significantly now we often use reinforcement learning to fine-tune our deep learning systems in the end to help different kind of you know loops that look at a global a global score that you wouldn't get at each word that you're trying to predict but you have for instance your scores in summarization or blue scores in translation and you can use reinforcement learning to look at the entire text and then update the deep learning system and deep learning systems aren't also just like now like layer layer layer of the same kind as you see in computer vision still a lot but they have no memory components if pointer components pointers are kind of underappreciated and deep learning for NLP and I'm sorry if I'm boring people who are like that deep into it but like pointers are allow us to predict words we've never seen a training time so the fact that in squad for instance in question-answering you can give answers based on a question that you've never seen before about the topic you've never seen before and it still kind of gives you some some right answers and generates words that it's never had seen in a training data it's pretty incredible those are like kind of underappreciated advances in the last couple of years and now what we're doing is just in our group we have this deck NLP data set and and first model where we have a single model that you can ask any kind of question and it still gives you a pretty accurate answer you can ask one model what's the translation of the sentence into German who is the president of Zimbabwe what's the summary of this document what is the sentiment of this tweet like in you think this goes on there's even one fun interesting example which is like what's the translation of this question into a sequel query and it'll learns to generate database queries and it's all trained in a single model so that you can then combine different things and it can even do a zero shot classification where you can ask it a question that you've never asked before a training time like Richard gave a keynote and nobody clapped is Richard happy or sad and like it learns to like point to the word sad based on that new context even though it's never classified stories before it didn't have a large amount of training data of happy and sad stories and it just kind of figured out the core current statistics so long story short we're still a lot going on the field of deep learning is growing and I'm super excited for the future thank you very much I think this is all the time I have for questions we get to get in already for the next panel but thank you so much you're yearning for gravity please stay where you are we're going to have the panel now on dating Union for and the research is going to be on this panel for you guys just coming in you guys miss the secrets of deep learning a very sorry for you but you can catch them on the video which will have published so let's give a big round of what we're thinking so the other you