Devreal

Text By the Bay 2015: Science Panel

Text By the Bay 2015: Science Panel

Recording: Text By the Bay 2015: Science Panel

uh that was an interesting day um full of talks from what I've seen floating around uh we had a great diversity engagement and uh uh it seems like we we hit some interesting context here so really appreciate all the speakers um and all the attendees and all the sponsors of course um I should mention that bis technology and lexalytics are our sponsors as well as nitro who is hosting this they're great Partners they're in the gallery and uh a lot of you had uh time to engage with them uh they do some really interesting stuff and uh David uh presented Seth has a talk and uh uh ol will be on the panel tomorrow so we have two panels uh today we have a science panel so we have questions of what kind of algorithms what kind of findings we have in NLP how ready is this for production and uh what are the problems uh people perceive and uh in in the field uh tomorrow we have a panel uh roughly called business of Tex the question is since this is an applied conference what what are the aspects of selling application using LP how how do we approach this market and and I think these topics are related and one cannot be done without the other so topic today is is really kind of uh scientific experiences and uh the plan is as follows uh we'll do the introductions which will take a few minutes uh I would say you know 3 to five minutes each uh and uh when you guys are introducing yourselves uh explain the most interesting problem briefly in challenges that you see for uh and will be to really work right as scientist what do you need to do to make it work and uh then uh essentially uh I'll follow up on this and uh we'll have a conversation and uh I think kind of two-thirds uh uh into our hour will open up for questions uh to the panel which you can direct either at uh the panel then somebody will answer or individual panelist heis um I'm the chief data scientist at banjo prior to this I fortunate to work with data science in proteomics genomics uh Health Care uh disease progression uh Sports strategies um amongst a few other things uh signal processing uh social discovery uh and now I'm working on event detections um banjo uh what we do is basically let you look at the world through everybody's eyes so we could look right now at this conference and see in real time what everyone here is posting about this conference independent of what hashtags they use and independent of where they post it whether it be Twitter or wayo or V contct or Instagram so obviously the main product of this is what we call a crystal ball which lets you look at the world and understand what's happening where why you know how uh in an instant uh the main driver behind that is event detection which then RS into text mining obviously uh natural processing and image classification um the biggest challenges are understanding context so we detect events based on the first post if you wait for a signal with volume you're too late so the single first post you have to understand the we detected an Amazon fire uh Amazon data center fire the other day a couple months ago the post said crazy day at work I'm sure it was a crazy day at work but obviously it doesn't have the word fire in it or concert from the band Arcade Fire it's not an arcade on fire so understanding that context both uh from text and and from Vision one small example from vision is we detect floods yet we don't alert Venice every single minute that there's a flood but when you don't just train on the data and say this is an image of a flood but when you train with context and you have a deep Learning System that understand context two months ago there was a flood in Venice and our image classification system was able to detect a flood in Venice which obviously looks very different from a flood anywhere else in the world um so I think uh that still poses quite a few challenges and my passion is not necessarily uh doing something that would be super enticing academically but enticing to a business perspective right A lot of times it's a loow hanging fruit that you can find a solution somewhere that gets you that product that nobody has and and it's not the same path uh you know an academic might take thank you uh hi my name's Jeremy Howard uh I recently founded a new company called antic um antic is trying to change how medicine is practiced um we think of medicine as answering two very simple questions one is what's wrong with you and the other is how do we get you better um we think both of those questions can be solved using data analysis in fact doctors solve them using data an Anis every day they just don't know that they are um the kinds of data that they generally use are um image data um and natural language data and uh to a smaller extent structured data so um I was previously president of kagle which is a machine learning competition platform before that I ran a couple of other companies um for the last 20 or so years I've been working with machine learning but in the last year or two I've decided that um everything other than deep learning is basic basically a waste of time and I'm not spending any time on anything else anymore uh I will say that in NLP you know deep learning is only just getting to the point that it's kind of stateof the-art but it is more or less there now in computer vision that was the case from about 2011 um and over the next couple of years it'll be the case in every other subfield as well so um that we spend all of our time pretty much at The Cutting Edge of academic research trying to figure out how to do things like put a 150 megabyte three-dimensional MRI into a deep Learning System and figure out whether somebody has cancer or not um using stuff like this we think we can you know potentially save millions of lives and make billions of dollars so I think it's a really exciting area thank you hi I'm Ben pedrick I work at judicata which is a startup around here that's making legal research sof Ware um basically that means it's a search engine for lawyers uh the reason NLP is interesting to us is that the outcomes of court cases are written by judges in plain text so if you want to know anything factual about what happens in a case you need to go look into the text and figure out uh when the judge says this person's sued over this issue and extract that information to provide it um the other thing that we are using NLP for a little bit and we be more in the future is we want to create a full map of legal Concepts and how they relate to each other and again pulling this out of court documents and case text and say okay if this is uh if it is unlawful to do this thing then maybe there's like certain evidence that you need to provide um and being able to extract all that in an automated way to basically generate a whole view of the law um my background before this is in more systems and not so much in academic or NLP stuff so I'm sure you all know a lot more about what's the state-ofthe-art right now and sort of what the big problems are going forward but I might be able to write a little more information on uh the problems that we face and sort of what the issues are that most face us right now um I think the biggest thing is that our requirements are very high in terms of precision and recall uh the body of uh doents that we need to look at isn't that big in California there's 180,000 cases that we care about which is really not that much um but for the information that we're taking out of it we want to be on the order of like 99% plus precision and recall um and so that leads us to do a lot of tricks to work around limitations of the systems that are out there thank you uh so I think I'll POS um one question to all of you I think there is one question which covers all the use cases so uh in order for NLP AI machine learning General uh to to work to apply in practice and I know Jeremy is done the faculty of Singularity University so in order for Singularity to actually happen uh right we need the uh systems to really understand our intent right and I think the key problem with C references right you say something and it gives you something else so uh how do we uh really infer the intent behind the the input right and I think like for instance in case of endorsements you know a friend of mine endorsed me for sandwiches because I bought them good sandwiches and what they really meant they wanted to thank me for buying them a sandwich uh right I mean probably this is very hard to infer but you know some sometimes people endorse me uh uh for Ruby or Python and what really they really meant that they actually visit my profile and basically what what I want to infer from this they have no clue what I'm actually using uh which is scholar in none of these languages uh anymore and uh right so so I can actually infer that they actually wanted to visit me and check me out right and and you know they provideed useful information to linkedin.com that they did so uh and uh I mean in case of social media obviously you can tweet anything and mean different things you can say oh great right like this SC train just went by and you need to infer that this is sarcasm I mean in case of medicine it's a little bit different right we need to infer what's really happening so the I don't know if there is human intent there is nature intent so it's probably uh um I can kind of quite uh put it this way but obviously in case of legal cases right the lawyer is looking for specific uh precedence related to the case so their intent is very clear they don't want to lose the case right and they want to find all the things which apply to them and good for them but not the other guy and they're mortally afraid of losing this one piece which the other side will use against them right and so they they cannot have anything but 100% recall right so so we kind of we'll know this if we know we use the use case of individual uh application and you know as humans we can very quickly understand so but what if you can talk a little bit to to your systems uh what do you actually do in in practice and how how good are you at in inferring human intent which may be different from what kind of The Superficial examination of the input will tell you you feel free to go any order I'll I'll jump in because you uh you jumped in on the endorsements example uh so what what I would say uh for the specific example you gave I actually would want to talk about another example because I think it captures the intent problem a little bit more um for for endorsements that's a very structured system meaning uh in a way it's it's guard rails for people to tag each other right and the idea was to build kind of like a page rank like graph of uh topics in the same way that Google uh looks at anchor links on web pages and how they link to each other and what the anchor text is uh we wanted to have a similar similar structured Corpus of linked entities across user profiles on on LinkedIn um and it had a semic constrained vocabulary so a lot of those endorsements that you're mentioning um some of them may have been organically entered by people saying okay I think he really knows Ruby but the it sounds like you're talking about examples where somebody a suggestion was shown to them hey do you want to endorse him for Ruby and they said yes right so in that case the system is inserting itself um so I think it's that example is actually less about user intent so their intent may have been oh I want yes I like Alexi and I want to show that I like him and so I'll say yes and they may not really know whether you know it or not um so I think that roach gets into this the the issue there is more about UI so I think it's more about um you know the the I think often what I see in data science and and text and NLP in in the applied world is a reluctance on on scientists or researcher uh uh on the side of scientists or researchers to be more proactive in the construction of the UI in the application and and to mold that in a way where they get good results or they get good data and there's often a trade-off between whatever primary metrics that are being are being measured in the system and and collecting good data so that that's what I would say that is more about what the challenges of intent I think another great example would be autocomplete uh and something like iOS um and so uh as an example so I think I turned off the actual uh maybe learning side with a person personalization feature of iOS in my in my case so I get worse results probably than normal but if you're just using that default normal system uh uh you'll get some pretty strange results which if it knew the context of the user you'd get better results right um how many how many people have weird autocomplete results in their ISS right uh okay well maybe it's just me um um so everybody else uses Swift K now everyone else uses Swift K so I'm be I'm behind but but that's strange right you shouldn't have to install a special application to do that um and I think to me doesn't have very good MLS people yeah and I don't know if it's about having good people or not but I think it does get to this idea of a mentality of using more signals and more data uh will get you closer to that intent I think um if it's just something that's running um online on the client side on the phone that's obviously got it very you know whatever they shipped as a a pre-trained packet is all maybe that it could leverage I think on that one you can just do it with better algorithms which I think swift key is shown and I think you can go further still right I mean the problem with NLP traditionally is it's been based on like engrams and bag of words and stuff like that and you know you use the word context which is exactly the right way to think of it context in language is nword long you know and who knows what n is so we need to use sequence-based algorithms so um recurrent neural networks are you know the thing that is finally allowed us to have language models that are actually accurate so soon as start people start using RNN for autocomplete you'll not have to switch it off anymore they work incredibly reliably but you have a good Benchmark and Google arguably does it much much better right so we know that this is more obtainable right yeah I mean they only use iron in parts of their voice recognition but I don't think they're still using it I don't think anybody's using it yet for the language model as I say in NLP deep learning is just in Academia at the point where it's starting to become the state of the art so it's still not really in anybody's products yet at that level but I mean to to to to me your question about intent is kind of I I would express it differently which is I don't think we nearly need to we really need to care about int or any other underlying cause or driver to me the question is simply can we do an action which is the best action at that point in time um that's that's all you ever need from a model right so you asked about medicine in medicine do I recommend the treatment that's got the highest probability of making you better at the lowest cost in the appropriate amount of time um uh and you know the same is true of uh you know if LinkedIn was saying hey here's a skill somebody has then it's actually okay if if we recruited that person would we end up liking them to that skill or in complet suggest this word do they end up clicking on this word um and this can be done through partly by being careful about the experimental design that you use for the data collection and partly about using the right kind of algorithms you know so rather than just using purely predictive modeling based algorithms you know making sure that you pick a an an objective function where the metric the metric you're optimizing is the of the uptick in the actual thing that you're the action that you're trying to drive rather than just the prediction that you're trying to make and then questions like intent kind of become you know they just fall out in the wash naturally I think I think it's a it's a good point you brought the recurrent neural networks obviously uh and and my line of work using bag of words is completely in feasible so we had to re uh but actually obviously those run into the problem of a shortterm memory that they have and then there's the more advanced Solutions of the ltsm and uh but uh I think even with those we've seen that even if you have the perfect understanding of language with more complex longer term memory uh neural networks for deep learning I think and and my problem space is a little different because somebody could say something that even to a human being if you look at the text you're like yes this is an event but when Harrison forward his plane crashed everybody in the world posted that same post as if they were there and clearly you can't predict that there's you know a plan crash everywhere on the planet and that's where I think at least in my field the context the intent comes surrounded by what's going on around it where if we are not only looking at the language but we're also looking at the images and we're looking at you know a view from space to know exactly what's in that location and we're looking at a his historical of understanding what that place is what time of the day it is usually on Monday nights this is what it's like and then this is what the person is saying I mean it's just like if you're reading a book you know they say oh he's painting a picture before setting the scene even in a book it's still we still visualize it you you still think of it visually because without painting that picture you can't really understand the text so well this is why this this is another reason why deep learning is so important because deep learning like we have the same thing in medicine you know we have the the the the answers that the patient gave to the doctor's questions we have the doctor's notes we have the MRI scans blah blah blah and so the nice thing about deep learning is it converts everything into you know this distributed representation space so we can now create multimodal neural networks in which we can combine the images and the text and the structured data into a single mathematical you know space where we can run queries whatever so I mean we're starting to see a whole new kind of model of comput utation now appearing where people are solving quite complex problems you know using these kind of dis representations it's a whole different way of thinking about things so in the end these rnns you know to me they don't an RNN is not a tool that makes a prediction an RNN is a tool that converts a sentence into a vector into a distributed representation which I can then add to the distributed representation of the images and the instruct data and so forth which is yeah it's a super cool thing we can now just starting to be able to do yeah the the thing I I I would step back if I you know re reinterpreting your question again uh another example would be a common one is query intent for searches right so uh an example we'd see at LinkedIn is um or I experience this myself if I'm typing in um uh I don't know Scala right it's pretty clear I'm looking for somebody who knows Scala right I'm probably not looking for somebody named Scala it's it's possible somebody's named Scala um but um I think that unless you are putting the right inputs or the right the right examples into your system and you set up the problem correctly um it's you're it's not like the system is magically going to infer that um at the level of Technology we have right so you to get to what you're calling query intent I think it at this stage currently right now production systems you actually do need uh a bit more data scientist or human effort to to craft um feedback loops properly and gather the right training data um uh to set those systems up uh let me actually pick up on this feedback loops I think that we've seen um a few talks uh and a theme is emerging that uh human in the loop right that can be useful or more broadly Active Learning um I think uh there is a lot of work in language was previously done by analyzing news corpora and was inherently passive right basically you're listening you're analyzing post factum with social media uh apps like LinkedIn you're in an interactive environment right so I wonder how much of this mentality is kind of inertia where we assume that people are doing something and we are just there to uh to be there and interpreted how much of uh actual experimentations happen for instance uh I think in in marketing it would be you know very efficient to just send a lot of probes right and and do a testing on live audiences and segment and and see what works instead of sort of you know doing a campaign the traditional way and analyzing it um so uh and uh and again I think I would I I would try to kind of c a bit more with uh jerem with the with the intent question because I think the medicine domain uh I think the intent is is pretty clear you know the cancer wants to kill you the disease wants to kill you so then intent is is is you can take the best action because um because you assume that the the you know exactly uh what the intent of the disease is right but in in in kind of in human situations we may have somebody for inance social media the fact that they're using social media doesn't tell us much they maybe they're browsing they maybe they're looking for something similar for L LinkedIn somebody may be a recruiter somebody may be decorating their profile somebody may be browsing their colleagues or checking things out right so so I wonder uh and and so the best action I think it's very clear what the best action is in in in the context of in the medical context right uh but in in terms of human context I think it might depend on your application because you may choose to engage the user through some kind of active probing right you can you can show them something which you would not show to other people right you can ask them to do something and I wonder how much of this you see as like for social media how much of this uh is happening how much of this can be happening and what kind of tools do we have to to do this experimentation in addition to to learning passively about what is it users are trying to do what I would say is I I think probably most of the systems that we work on have some user or customer at some level and I think you know uh if the targeted prediction problem doesn't seem it doesn't seem like user intent applies maybe it it I'm pretty sure it'll apply in whatever interface is uh being presented to the user um I another example of this I think I see is uh uh systems often neglect uh behavioral data right so even if you're mining across all of your uh user data or tracking data or marketing data to see um you know try to segment or predict what the intent was of a particular query um you got much better results if you can personalize it based on your behavioral history um so an example I'll pick on is anyone from Yelp here no um I think somebody from Open Table is here so uh I won't pick on open table but uh I'll pick on Yelp in that I'm a big Yelp user um but I feel like for example when I search on Yelp I have to uh put in the same search facets over every time open now near me like I have a certain search behavior that I use um but it feels like it doesn't remember it right the next time I use it um and so that's something uh you can call that another type of intent which is we we tend to do things in a h in a habitual way um and if you can learn those habits uh it makes things easier it's it's mention y because I I agree that you know the their actual experience with app seems that it doesn't know much about your history however their batch mode seems much better because they send you all these emails which are taret like I live in is Bay and they try to tell me here what's happening in ISB they have all this cute subjects which are supposedly will make me click and open this email right and so I think maybe you know that part is is just in a different group right but I think there is some part of this experience but it's not it's not in the moment y like y like the anti- machine Learning Company you know they just there's so much more than I mean they could be doing like that's why you know like you shouldn't have to manually remember every time in I mean just the most obvious simple machine learning would be able to figure out pretty quickly that P looks for things that are open now or prefers things that you know that are open yeah yeah yeah or or or or whatever this is all this manual programming can be massively cut down by using just most simple machine learning models that's a company that chooses not to for whatever reason I think a lot of companies are like that though probably most probably yeah because I think a lot of CEOs don't understand appreci iate or trust machine learning so direct the engineers to use heuristics all the time so the other thing that we've found is that um for some of this for generating intent it's not a static process it can be an interactive process and so in terms of generating query generating Yelp uh something like that like even just being able to suggest stuff out of your history and not set it would not be that hard for you know to present hey I want to look in like an auto suest sense I want to look for this thing in things that are open now or in things that are uh care to I don't know my dietary concerns or whatever what you said about the CEOs I do see that a lot and I think um it it's a problem on both sides a little bit I I think a lot of times it's the miscommunication where uh you might have a team of data scientists that is not very business oriented and so if their goals are are not you know completely digested into a data science problem then there's this miscommunication where you can have a data science team building great models but to the business point of view they're useless and so you need someone to bridge that gap of okay stop thinking about data science for just a second what is the business goal now transform it into digestible bits and and that it's unfortunate the data scientists involved in like talking to customers and designing products and doing engineering places where the data science team sits out there in the sun tower they just create kind of academically interesting but ultimately useless kind of things Jeremy what do you think um there are some in some sense like for things like social networks like Twitter Facebook LinkedIn maybe it's easier because the people building these systems are also using them maybe you know heavily um it it seems like it would be harder uh in a company where uh Yelp is maybe again counter not a good example for my point but um if you're not if you're using an application that's being sold to a completely different type of user um it may be harder to for scientists to connect to the user like uh I don't know what for sales people say uh if you're a data I mean I I absolutely think there's no reason I mean maybe medicine is a good example Med I was going to say like not many of our team have long backgr in Radiology but we still spend most of our time sitting next to Radiologists looking at MRIs trying to diagnose disease uh and so in the end the Radiologists at our company get reasonably good at machine learning and the Machine learning people in the company get reasonably good at Radiology um and I think that's the the trick to avoid the problem that you know Pedro was talking about where there's this kind of mismatch in the end when the two people sit next to each other it's kind of rather than um pair programming we're kind of like doing like pair medical Diagnostics and it's like okay does this guy have lung cancer or not and you have a radiologist and a deep learning practitioner sitting next to each other interactively trying to figure that out together is that something in the legal domain yeah we have a very similar process where we have a handful of people who are lawyers and handful of people are Engineers if we want to extract some kind of data out of a case we'll sit down and look at it all together and we both get a really good understanding of first like what the information is and how it's presented and how we can get it and then also what the limitations are in terms of the uh the automated systems that we can get it out and where the hiccups are going to be i' like J or an analytic or interesting examples right because like lawyers and doctors of things that probably all of us were told we were meant to be when we were growing up I don't know I know why I was and like here we are kind of like in the process of of kind of automating and disrupting these industries with machine learning doctor law I want to pick up actually on this interesting thing so so we talked about the dichotomy of data scientist versus business right and so the I think this is a consensus that data scientist should be closer to the business goal and this can be achieved in different ways for instance uh so at nro actually I'd like to you know propose this and hope we can do that that I want to be a sales person for the day right we have a sales team and you know I'm personally like I want to make a sale right I want to talk to the customers I want to understand what they're doing and and I want to convince them to buy a Nitro subscription at the same time I want understand what is it we need to add to it as a cloud product because we can just push things into the cloud platform right so so I think that that's that's clearly beneficial the one thing which uh I don't know if we have a consensus on is how much uh data scientists should be uh Engineers right so we have I think there are two models in one model we have kind of a primon model where the data Sciences can use whatever tools they want they can use R and they can use MLB and somebody else have to take the model and redo it in a production language such as Java scholar and the other model is that everybody should be able to implement their own ideas right it's not reasonable for somebody to implement somebody else's ideas if they have their own so um um and and it is possible to uh to hire people who can do both things right essentially to develop algorithms in in the language or in a setup which is either immediately can be put into production or just need some expert performance improvements if that's the case right but and and that kind of increase the iteration cycle but I wonder from your experience how how feasible do you think uh that goal is that at this point you know data scientists should be uh Engineers or are are they in currently different species uh this is something I've talked to a lot of people about and thought a lot about um so I don't have like a very definitive answer at this point um I thought I used to think I had a definitive answer but I've I've kind of softened over time um what did you used to think I used to think uh the everybody should try no let me let me take it back um I do think everyone should aspire to being that at that intersection right having uh solid engineering skills and uh some research skills and understand the business right I think the reality is the the reality is that it's it's anytime you're doing those kinds of intersection operations you're going to reduce the set of people that you can possibly hire right um that's a smaller population so the the pragmatic reality is that um but not necessarily because the other thing you can do is you can hire PE people that have this strength but then you try to teach them the other things so no no that and that that I think is something I mean I spent a lot of time both you know hiring and thinking about this at LinkedIn as well as working on it in a meta way at LinkedIn when we looked at the skills Gap and the economic graft and a big part of our mission was creating Economic Opportunity and I think a lot of the analysis that you see out there in the media around the skills Gap kind of misses the point when they say there's a shortage of X or Y um I think our part of the problem is actually in the last you know 10 to 15 years um the recruiting mindset has uh leaned on tools that had very simplified ways of screening which are like very just keyword matching based right so we have a datab Bas of resumés we have um a job description and they do some pre-filtering right and if you don't overlap uh on the keywords maybe they pre-screen you out but that omits what you just said which is you can hire and train people and you can invest and training yeah say high after what somebody can do for what they have done yeah so I I think the me the mentality that I take now to get the short circuit my answer here uh uh is uh look for aptitude and hire for aptitude and then level those people up um but uh in terms of what tools do you use uh do you just let once people are in the door do you do you try to level them up in engineering or do you just let them use whatever and throw it over the wall um I would lean towards uh trying to train people up um uh but I've seen I've seen companies doing both approaches um and you have to pick what works through your own environment I think what we saw tle was that most weally work with big companies weally saw that systematically on the side of creating everything as an engineering project you know and like so went allowed to use r or whatever the hell they want to use and when you say they you mean our clients they yeah so kaggle's clients you know so the data scientists wouldn't be allowed to use like whatever tools you know they were told by it that these are the tools they're allowed to use and it wasn't just programming languages but at least implicitly it's like oh you know if you use anything more than logistic aggression then Management's going to tell you you're wasting your time being a primadonna or whatever but then you see what happened when they actually buil something along to kagle and they're bring like like all state you know they're bring like a critical question like how do we figure out who's going to crash their car a cost of soci load of money it's like the most important thing an insurance company can do and uh the answer basically is well use a gradient boosting machine instead of logistic aggression and you get an improvement of 270% in accuracy and you know all state absolutely had the people that could have told them that there already but they just never gave them that freedom so I would say you know my view definitely is go the Primadonna approach you know like find out what what can be done but then have them work in a project to implement something that has the kind of latency characteristics and fits into the infrastructure and so forth knowing that okay 270% Improvement is possible we can get 240% Improvement if we spend three months on it and have this latency and blah blah blah you know I definit I think that's letting data scientists use very interactive rapid prototyping tools is very valuable the Stu where answering that question well is critical to your bottom line to his point then would you say those data scientists should then learn uh if if the standard is to do it in Java and schola should they be learning that as well uh I yes absolutely but they shouldn't be using those tools to find the answer they should know enough with those tools that once they know what can be done that they can then work you know P program with you know low latency High scalability Scala Engineers to kind of sit next to them and work together so then when the Scala low latency High scalability guy goes oh we have to make a choice here like we could memorize this piece there or we could recalculate it and the Machine oh you could actually do neither with just a 1% decrease in error we could avoid calc this all together because frankly this is kind of a bit of an unnecessary tweak so that's why I would kind of say yeah have people that have specializations but also try to have everybody know a little bit about what each other are doing and work really closely together kind of like the lawyers and the engineers working next to each other yeah I agree it's bridging that same Gap um but also it's much easier to have an ideal solution and then decide if you have to cut things and minimize the decrease in performance accuracy then the other way around you can't do it the other way around come up with a suboptimal solution and engineer it enough so that it's that's never going to go that way right so you have to start with what's the best and then start making the choices of the engineers I want to ask uh Ben about uh this his specific case right so I think um so jiat has this very interesting domain where for instance you need to do legal search but you cannot really use Google or something like that because you you have very specific constraints right you have to have a recall so so you know on the one hand we have a lot of different kind of um machine learn problems which we can apply standard algorithms to but in judicata case and I think in many other uh companies which have Corner cases right kind of represented here by judicata you have really uh specific constraints which make make it not uh feasible to just take off the shelf Solutions so you have to simultaneously kind of engineer on algorithms engineer on data pipelines and then you have your own uh human lawyer reviewers so I wonder if you can kind of uh illustrate on your use case how all these things are coming together and how much you know of of the Shelf Solutions you can take what kind of U algorithmic differences improvements uh you need to do to make it work in your business your business case mhm so uh so one thing I'll mention so for us most of our data already exists because it's core cases that have already happened um so we can ask humans to go look at it and handle the cases that the algorithms don't get so it's kind of this scale of at the beginning or with no automation you could ask people to do everything but that's really expensive so you automate as much as you can and then leave the rest of people and it's more about finding the one that's the best cost for us um and so for us finding there's like a limit to how far we want to invest into an automated solution and we kind of want it to be tunable on the scale and we have a small domain so we can take off the shelf stuff and that gets us like pretty far and then we can write some our own stuff on top of it that gets us a little bit farther and then we hit a point where we realize that getting that last little bit is just way too costly and so then we can stop um and so as a result we're I think more due to the uh how limited our do main is we end up doing a much more rule based approach which I'm sure makes all the machine learning folks cringe a little bit but um but we're able to push that pretty far in a very short amount of time um which is really effective for us so we kind of have the opposite problem with you guys where we are more Engineers who are like dabbling in data science and data scientists who are dabbling in engineering um but again we can get it to where we need it to be let's that's a good balance right I so I think I think I think this balance can also change by by application I think this is a good time to maybe uh have the audience ask us some questions uh so uh if you want to ask a question please come here and use this handheld I'll just give to you so the we have the recording uh and you can address it to the whole panel or you can address it to individual folks if you have questions oh yeah please and if you want uh we can line up here and so we can be efficient hather I found that the need to improve on our solutions to make a human believe that we're doing a good job to uh put a lot of domain knowledge into the system so we get the right terminology so the little nuances are discovered uh how important do you see that and when do you think it's good enough to uh to stop with the terminology work uh so also I suppose i' there would be medicine ones and and legal ones that you you can't quite borrow you have to sort of uh tailor yourself can you say what what application or what so ours is in consumer product so there's just a lot of different ways to talk about a consumer product and what they are so for us we build lexical ontology to help navigate the space of consumer products there isn't one out there Wikipedia is not um able to help us there and when we use like a a distributional model to get the multi-dimensional vector space so you can tell that it's making some very very simple mistakes because it's not getting the nuances of of what people are talking about and then our humans um that are looking at what we do can tell us that we're doing it wrong because there's some obvious things that you're you're missing out that this is not a color this is not a fabric type it just happens to be sort of similar is this a B2B product Arch is an online advertising so we uh help put insert insert links into um blogs or user generated content and put it to affiliate marketing um so we found to be useful but I haven't heard whether other domains have benefited from that and it's unclear to me when you stop and when what when you start yeah I I mean the the example I gave at the beginning of uh uh skills and endorsements that was kind of the entire premise of start that project was uh give an example I don't remember the exact numbers but uh we mapped those entities to Wikipedia as part of uh the pipeline but it it was more um as a way to we use that as a signal in uh some of the disambiguation um but we didn't strictly uh adhere to Wikipedia um and I think only a fraction of the skills or uh areas of expertise had a corresponding Wikipedia page or entry um so the vast majority of those skills are actually uh new entities that were defined on LinkedIn um and the other thing with uh something like Wikipedia or prebase is probably richer but um they grow out of these often grow out of communities where the rules are different right so what what qualifies something to be an official say Wikipedia page it has a bunch of strange Community rules around it around notability and things like that um so you always have to consider the source of those those data sets or the label data that you're getting um so I think that's one of the big one of the most powerful things you can do in a domain specific application is uh invest some um uh some time and energy into uh and Google Google does this as well and they just rolled out what a Knowledge Graph they have medical you type in symptoms you know I I was sick last week so I typed in you know whatever my symptoms were and and a little uh Knowledge Graph thing pops up on the right with uh Google's inference of what what disease you have or right um so but they they hired I think armies of contractors combined with uh you know machine learning to build out those domain by domain um so I think that's I don't know where we head you know 10 years from now but it feels like right now that is uh the plan of attack that a lot of people are using in Industry whether it's Facebook or uh Google um or or LinkedIn or entity classification or whatever is saying a concept has a taxonomy or it has a category or whatever and I just don't think that's accurate it's not nuanced enough and a concept an entity actually has a nuanced multi-dimensional kind of but if you do same things and and and and as a dis representation you know you get to express it in 600 Dimensions you know and those 600 Dimensions can interact with other representations in complex ways so for example I mean I don't know sort of worked on LinkedIn because on LinkedIn you didn't have a lot of prows but like on on big link I totally think it would because you guys are looking at you know pros and so you could really understand a word quite deeply in that rich multi-dimensional context and then find the appropriate Subspace you know based on that I I I I'm not disagreeing I just I think that you know when it comes to features that that totally makes sense but when people want a single URL to click on and so WebMD it may be okay hand foot and mouth disease or something right so people do want to label so I think it's not so I don't disagree like so linear like yes using deep learning maybe a better the the right uh approach behind the scenes um I'm just advocating that getting rich labeled data uh or generating your own Rich data um uh from your audience from your customers if you can uh or hiring so Google the way I do it is they hire our contractors uh uh essentially to like start building out some some of those domains um but they can still use deep learning along with yeah so I guess what I'm saying is I feel like more of that should be done through through reading through through through through Reading using deep learning word to type stuff and that kind of gives you those contexts and nuances and and then where you need to have human labeling like you do at judicat and like we do an lytic that's where I would bring in active learn you know I would I would use active learning to then have this exp if human labels label the highest value points I think that brings up to the point of tradeoff and loss of a business decision that could be accomplished within a week or two or weeks of work versus a machine learning approach that would take a few more months to get there and I think the fear is a lot of companies make the decision of okay let's wait four months to get it right without having to use an easier path and that's one way obviously in business you want want to win and you want to do it quicker the problem is when they take that other path of perhaps taking a momentary easier solution it it it comes at a loss of doing it in a way that it hampers the vector representations because they embed the hackness of this data into here and then this can never grow anymore so it's just finding a balance of making sure that this is not harmed and you can still have that in a couple months but in the meantime attack the easier problem yeah and I I I don't know that the consultant or cont contractor approach of humans is any any faster or easier it's actually probably way more expensive um in the long run I think I think for your average CEO at first it seems more appealing like you know a kind of fistic rules based manual labeling approach there's a clear path the steps whatever El to me I think the machine learning approach often is faster and cheaper but it's just somebody who's it seems high seems a little mysterious something so so if the company hasn't done what you're describing before that that is almost it's not infinite risk but in you know the perception would be it's very high risk okay I'm curious you mentioned Vector representation so I'm guessing you're using something like B to VC or some kind of vector yeah we've used VL yeah were you I'm just curious were you looking at direction information too or just proximity like overall distance measure uh just yeah basic SK skip grams of uh of words um on on our text and then um do some exploration of the space finding clusters in there to make sure that when you say eggplant it's the color versus the vegetable or when you say cancer it's the zodiac sign versus the disease there's a lot of places where it gets really convoluted so have you used are you because there's you can actually break up the same word into multiple meanings and then have them represented in different places in were to vac right right is exactly so I I don't know if you tried that but sure there's really interesting stuff there it's amazing but uh it's it's it's losing the edge for a human a human can spot us and tell us wow you're uh not not bad but not good enough uh we need to layer in a lot of human knowledge to get it above the bar to say ah okay now you're in now I can trust you that's what we're finding so it doesn't sound like you for example Jeremy that you're building a terminology um terminology dictionary you know in medicine where it's there the distributed representations of words that you see in medical reports and I mean in some ways we have it a little bit easier because the Lexicon and structure of say a radiology report is a little more well defined than some for me I feel like what we're seeing Pi this simplii maybe represent possible in a more General sense in to time you know like I think it's kind of if you see some fracture for example in a radiology report it's very unlikely they're talking about uh cracking for oil you know they're probably talking about a bone for instance so you I think we're working in a somewhat easier domain it seems to work pretty well for us we do spend a lot of time um just manually looking at words manually curating list generating it quickly and just going through and checking it off um and again I think it's because our domain is small and we have a a really good sense of the bounds of it which I guess doesn't really help your question because your question is kind of what is the bounds but for us since we know what we look at um because it's search Bas we can just kind of look at user queries and say okay these are what they're looking for these are the things that we need to understand that maybe we don't right now that helps us a lot a lot inform uh how far we need to take it and let's us round out our data pretty quickly uh do we have any other questions from the audience uh if not maybe okay we have one awesome thank you so what are your guys' thoughts on using the representations as a way to augment something that's labeled so in this case like using wordnet as part of the word as part of the neural word embeddings for example so you can get you know automatic disambiguation with augmented with the augmented senses and then also like in the domain of factoid QA like something like Watson where it disambiguates the answer for you gives you one answer and it's focused on giving you one answer and you use the mixed representations as part of the overall curated list or like possible set of answers so basically using that as a signal for answer candidates and the like yeah I think I think the general idea is good one I to me the the trick to it depends kind of on how much data you have right I mean if you have like when Google actually did word to they had they would reading books and they have Google books and they kind of had enough data and I don't I don't think things like word net were going to add a hell of lot to that for us you know we have x 100,000 or x million medical reports you know uh or images and anytime you have any kind of data shortage like that the trick in deep learning as you well know is um data augmentation and so to me things like using worknet and stuff like that are ways that we can do in NLP the kind of data augmentation that images we do with like Reflections and rotations and blah blah blah it's kind of like hey let's generate some uh semantically equivalent grammatically correct sentences using external sources like I think that would be how maybe I would use word net but I I'm not sure that anybody's you might know has actually studied whether that compared to using word net for postprocessing for example is a better approach um but for me to me you know like the the right way to do deep learning is augmentation you know that's that's my bias yeah I I'd add in um you know so I'm no longer at LinkedIn uh now so uh the the projects the things that I'm working on now you're starting you know it's one thing when you're at Google or you're at uh LinkedIn and you're starting from a large data set of you know hundreds of millions or or billions of of of documents or users um but without that I think when you're bootstrapping a system I I think go for it it you know it pays to be a little Scrappy um the the thing that you know it's a nice problem to have if you know this whatever system You're Building gets traction is useful and then you have to replace that stuff um but you may hit a point you probably will likely hit a point if it's successful where you have to re-evaluate that and I think in the example of skills um we bootstrapped there was a old Pro section on LinkedIn called Specialties which was basically a pre text field that people could put um comma separated uh keywords or they could put whatever they wanted so some of them were sentences some of them were comma separated Fields um bullets things like that and we started by mining unsupervised from that section um and using Wikipedia and using other data sets initially to build uh you know the set of topics um and then over time when people started adding skills to their profile we switched completely to that data source um so and got off of all all the other things uh so I I think uh like in general I think that would be that that would make sense um but you're probably on your own what he said uh any other questions from the audience uh if not maybe we can wrap up with kind of a question to every each of you and you can answer them in order uh what should we uh as a community of uh data scientists and Engineers work in LP do to make NLP more efficient in practice so from what from your use cases from your applications where did you see uh some areas where you wish you could invest more effort into learning about algorithms maybe throwing in more brain power into something what what is the kind of uh lwh hanging fruit what should we go after in in the current state right where would where can you make an best Improvement what do you think will get us closer to Singularity Singularity um you're assuming we want to do that okay at least you know make Siri not really like get really really confused right like get to a better Seri let's just make that kind of a uh better wish right um so I would say your current neural networks you know like if you're an Academia pick any state of the art problem and redo it with the current neural networks every time I've seen somebody do that in the last 12 months they've had a paper in a high citation journal and kick five at the same time yeah well exactly you know there's been plenty of recent um Journal articles which is basically hey I picked 12 challes problem the same obv I just surpassed every single one of them better combine it with reinforcement learning uh yeah normally that would be 12 papers but hey I don't have time to write 12 so here's one or do the same if you're an entrepreneur do the same thing in in startups like pick you know problem to solve x and do that plus learning it's like it really reminds me of the other what should I do should I do like e-commerce books uh stuff in government uh blah blah blah it's like X plus Internet it's like well every single one of them you know internet's going to impact everything it just pick anything like plus Internet it's the same I would go there like if you if you made another swift key clone but with an irnn I bet it would be so much better than everybody El you know keyboard program it would take over for instance to pick something at random pay day yeah I think for us um the the problems we're facing are not so exotic that all these approaches that everybody is working on but in health the biggest issue that we face is that because it's so domain specific the data sets that are out there don't work very well for us um making our expensive so much in previously mentioned getting data that is you can feed in that is useful to you big issue for working in a domain specific field okay well I mean now now feel so reenergized that about deep learning I want just to go back and like really dig into this thanks thanks to Jeremy uh but I think this is actually this is true because it to things like documents and I mean obviously the recent work from yon's Lab it's very encouraging although it does seem a bit magical right because like he just fit the sequence of characters and don't care how they make words and even even in which language that still looks like magic but maybe it will get us there thank you uh everybody this was a great great panel really appreciate you your contribution uh the plan for the rest of the evening is as follows we have a taco bar we have ice cream and wine reception and we have some board games so the idea that if you want to hang around and Sh to people we have board games laid out on the tables you know you feel free to play them or not but we'll have the tables and uh we did this at Scala by the bay it was a really great experience because we could just hang around and and chat you know um around this um uh games um or actually play them uh it's up to you we have a lot a lot of them and uh it will all be happening upstairs