Devreal

From text to knowledge via ML algorithms...

Event: Data by the Bay

data.bythebay.io: Xavier Amatriain, From text to knowledge via ML algorithms - the Quora answer

Recording: data.bythebay.io: Xavier Amatriain, From text to knowledge via ML algorithms - the Quora answer

okay uh thanks everyone for being here um it's a pleasure to um be presenting in this conference for the first time um maybe before we get started and just to get a sense of um who I'm I'm speaking to let's have a show of hands um how many of you in the audience are data scientists okay like that's most of the people how many of you are Engineers like machine learning Engineers or on the engineering okay uh so another important question how many of you are Cora users okay so I'm going to be talking about things that relate to kind of like the intersection of data science and engineering uh I'm not going to be talking about how we organize and we Define the those two teams and how they interact together although I've talked about that in many occasions and you can read some of my answers on Kora but um today I'm going to be focusing on how do we use machine learning at quora to get knowledge out of the text that we have so before I get into that it's probably good to understand how do we think of ourselves at quora and I think some people simplify quora to saying oh it's a Q&A site and that's for us I mean it's true that we based uh we are based on the Q&A Paradigm but really that's just a tactical issue what we really how we really think of ourselves is and this is our mission um a site that or a site in uh a company that really targeting to grow and share the world's knowledge and there's many ways that you can do that and you can think of maybe the closest example Wikipedia deciding to do that following the Paradigm of the encyclopedia right you can grow the world's knowledge and share it using the encyclopedic approach we think that the question and answer approach is actually better nowadays to grow and share the world's knowledge but the basic idea is we are in this uh world to basically grow the knowledge and to share it and we do have of course millions of questions and answers we do have millions of users and we have thousands of topics that uh we will use in different ways to share that knowledge and to grow that knowledge because of that there's three different aspect or Dimensions that are really essential to us and we care a lot in everything we do and especially on the algorithms that I'll be getting into later those three things are relevance quality and demand um relevance basically means we want to make sure that we get the right knowledge to the right people and that's really important because that's one thing that makes basically knowledge sharable and it also makes it grow right because we want to get the right question to the right person so that person can answer it we also want to get the right answer to the right people so they get interested and they generate more knowledge quality is really important and something that sets us apart from many other Q&A sites and I think that's an important uh thing that relates uh to our mission again because knowledge means by definition you have to have a notion of quality right knowledge without quality it's basically noise it's not knowledge it's just information that produces noise so we need to care about the quality of what's being written in our site and we'll talk about that and uh last thing is demand so we also care about what is the knowledge that most people want to know about right so you could write a lot of things about a very obscure um very narrow topic that nobody cares about and you're really wasting a lot of energies of the person who's actually generating that knowledge U however there might be some knowledge that a lot of people want to know about and nobody's writing about and that's really important because we care of course because of the also related to the relevance we care about sending the right knowledge to the right people we also care to understand what's in knowledge that people want to know about okay so then to summarize sort of like um what is our goal is given the text in a question or an answer we want to understand a number of things we want to understand the quality of the answer or the question we understand who will be interested in that how important it is um what is it about what it relates to all these things are going to matter in terms of optimizing those three dimensions that I talked about okay so the truth is we have a lot of text but we also have a lot of other data that we can combine with the text in our algorithms um so speaking of text and I'm going to go through this very quickly but all of the next slides you can actually refer to the data block on core and read more about it um we've this is an example of how much information we can uh find in the text and we can find in in our textual um data uh in this case uh when when a data scientist in our team was mapping the discussion over time looking at the text and looking basically at what are the relevant key Concepts how they relate to each other over time and you can see things like on this graph over here where you have a focus term like Obama and and there's different things like election Syria immigration and so on and how those terms over time change and how we can track topical interest in our population and this relates to demand but topical Demand right for example this is just an example it's a very long blog post that you can uh read if you're interested in uh how we extract Knowledge from data um but again there's a lot of information that it's not only text right if you actually look at our interface and this is uh an answer from U by Andrew in you'll see that there's a ton of stuff like there's the topic there's the question there's the user there's things related to the user there's the answer itself there actions on the answers things that users can do like up vote down vote comments all of that so we have text and all those things around text that we can infer a lot of information from so this is a graph that I usually use to explain explain the complexity of our data ecosystem um so basically we have questions and answers and that's the Q&A Paradigm but we have a lot of other stuff right so we have users obviously that interact with those questions and answer so users can do different actions they can up Vote or down vote they can want an answer they can sorry they can want a question they can Reas a question U they can answer uh a question and they we also have an overlay of social network so users can do things to other users like they can follow they can endorse on a topic and so on and speaking of topics we also have a topical Network that is overlaid on top of all of that because questions have topics but users also are endorsed on topics or write on topics right so basically we have kind of like three different overlays of data relations there's the textual one then there's the user one and then there's the topical one and all of those interact in this way and they're taken into account in the different algorithms um because of that we can study complex propagation effects in any of the networks so in this blog post Shankar was talking about how upvotes which is one of the actions just one of the actions that users can do on the answers actually propagating the network and again this is a very long post that you can refer to where you can see how the upvotes propagate on that overlay and how the different connections including the social ones affect the propagation effects on the upbs um on the topical overlay we can also study like how do topics relate how relevant they are what's the hierarchy that we have between the topics how different answers and questions and users relate through topics so this is another block post that you can find on our data block um and finally uh we can also look at usage patterns so this is not one of the main overlays that we have as I explained uh right now but just to give you another sent of other other things that we can look at um here in particular we were interested in understanding what are the usage patterns and what they tell us about people and this is a very interesting uh block ports where you'll find things that might seem obvious like uh what is what is the day of the week where people actually search for questions related to hangover or for going out or um being laid to school uh so there's like different things that obviously over time and different patterns of usage and that's also um an example of sort of like side information that we can use okay so I'm going to move now to talk about different machine learning applications um is there any question so far anything that I can clarify no okay I have one yeah so um uh someone post a question right why do you not cwl and find the answer if available on the web um okay so the question is why don't we extract information available already on the web uh and I I think that's uh it's a good question it's important to remember that the main part of our mission is to grow the world's knowledge right so uh that's very different from Google's Mission which is to organize the world's knowledge and by organizing mean whatever is already available on the web our goal is and our hypothesis is there's a lot of knowledge that's still not digitalized and still not on the web what we want to get is you know that knowledge that in your head and put it on the web so that's part of our mission and it's kind of different uh from going out there and grabbing stuff that is already available of course people in order to generate knowledge they might refer to things that are already on the web and they do that all the time right they synthesize different sources they put them together they answer a question but that's different they in doing so they are generating new knowledge they're not really just referring to something that already exists uh as a matter of fact we don't like answers that are just a link right it's like hey you can find this here well if you can find it here then you're not generating new knowledge do you extract information fromia no no okay so one of the dimensions that I was talking about was quality how do we understand quality and what does it mean to understand quality so I'm going to uh talk about this through a couple of examples one of them is answer ranking so answer ranking is an interesting example from a data perspective Ive and from from a machine learning perspective of how you can turn a real world problem into something that you can optimize through machine learning algorithm um the question we face and this is uh explained here in the goal is given a question and N answers that we have how do we rank them according to Quality and of course I think the the key question here is like what does quality mean right because you can Define quality in so many different ways fortunately if you ask people that work on the product um so that's what we did uh the machine learning engineers and data scientists talked to the product people including the CEO and said hey how do we Define a Kora answer a good Kora answer what does a quality answer in Kora mean um interestingly we have that in the in the site so we have a policy and we have an answer that says how do how do we identify a good quality a good quality answer on quora so when we read that definition we see things like truthful reusable it provides explan explanations it has references it's well formatted it doesn't have spelling mistakes so on so forth right so you can you can actually list all the qualities of all the properties of a high quality answer so what we did then is basically turn all those qualities into features that we optimize in a machine learning algorithm right so we have features that relate to the text quality itself we have interaction features that the users actually support or don't support that answer we have user features for example it's really important to understand what how exp how much the person that wrote the answer is an expert in that given topic or that given field and so on so the basic idea here is you don't have to um come up with a unique definition of quality you need to understand what are the different dimensions that play into that definition and then optimize your machine learning model to that yeah how do you deal with the bias because if you rank it higher because it's higher quality then you make it more up you're using up as one of the uh yeah I mean if you rank it higher sorry oh yeah uh I'll repeat the question the question is there's some interaction features that will um receive some form of feedback because if you rank the answers higher they'll get more interaction so they might get more up vat uh that's uh mostly true but it's not entirely true so the fact that you rank um an an answer higher it's going to generate more feedback from the user but that that feedback doesn't doesn't necessarily need to be positive that's one of them right so you it will most likely get more up votes but it's also subject to getting more down votes and other kind of reactions um and there's other things that actually the model has learned to Value more than some things and some of them even relate to the timeliness right so we also might take into account I'm not saying we do or we don't but we might take into account like how recently those up votes came in or they didn't come in and the reason for that obviously the best answer to a question might also depend on was that answer given 10 years ago or is that answer recent and how does that pan in so all these things are going to matter and actually the system is pretty Dynamic so you'll see the things come in and out and it's very likely that some things are pretty new they'll be ranked higher than things that were uh higher before yes so you give a numeric for quality that how is that done so there is there are features that I understand but then what is the target value for the yeah I mean the bottom line is you need to train like any machine learning model in this case this is a ranking model uh so it's a learning to rank approach so you need to give it some positive and negative examples or you need to give it some pairwise ranking examples right so that's that's the thing you need to generate that training data and there's different ways that you can do it you can do it either through like ratings that you give uh in the lab and experts can give that or you can also use the feedback from some users in the product or you can combine all these things the bottom line is you need to have that Target that training uh in order to generate this ranking but it's not any different than any other learning to rank approach okay I'm going to move on um other things things that we do in the context of understanding quality are related to the notion of moderation right because um obviously if you want to keep the quality High you need to do things like make sure that there's no spam there's no harassment there's uh no use of uh say bad language there's there's a uh a number of things that you need to understand and filter out of the site and and that's not a trivial problem right um I mean spam is probably the easiest one but all the other ones are kind of hard uh to understand so how do we do that um you might think ideally you you could hire a lot of people and have them all sort of like basically manually check that every question and every answer doesn't fall in any of those uh different categories right that is very expensive and doesn't scale so what we need to do is this hybrid model of of machine learning plus manual curation so how that works is through something we call the moderation cues so the idea is you have a machine learning model that let's take the case of spam because it's the easiest one to understand uh you have a machine learning model that decides this is Spam and just takes it away there's other things are clearly not spam so they're taken away or they they go into the site and then there's the gray area there are things that I don't know this could be spam I'm not sure the algorithm is never going to be 100% perfect so for those things in the G gray area we route them to the moderation queue we don't only route them in a kind of like first in first out cue we actually have a priority there are things that are more important to understand quickly if they're high quality or not because we anticipate they're going to be very important uh so we feed them into a moderation queue that has some notion of priority and then we have humans going into that moderation que and looking at the next thing they need to look and say oh yeah no this is Spam so remove it so that's the thing we do for again we have different cues with different kind of uh people that have specialized on different aspects like spam harassment and so on so forth um and we have the all the different machine learning algorithms that are tuned in order to feed into those cues is that how you manage your duplicate or similar question no I'll get into that later the question was about if that how we manage duplicate or similar questions okay um so the next thing is we've seen a couple of examples how we understand quality how do we understand relevance right and by relevance means like is this interesting for that person um typical example of this is feat when you go to your homepage at quora or you open the app on your mobile you're going to see your feed which is basically what we think is the best stuff in Kora for you and what does that mean I mean what's the goal the goal is actually to present the most interesting stories for a user at a given time and is interestingness is uh a loaded word also interestingness means topical relevance plus social relevance plus timeliness uh all that needs to play in uh and stories importantly means the combination of both answers and questions so we're trying to optimize the reader experience but remember our goal was to grow the world's knowledge so we also want to get those questions in there in your face if we feel like you can really answer those questions um so the idea here how do you do this the machine learning approach is a personalized learning to rank um there are many challenges that we need to face here like there's many thousands if not Millions sometimes of candidate stories we need to make sure that that's ranked in real time because we care a lot about timeliness as I said before and again we optimize for Relevant we want to make sure that the thing that is up there is the most interesting for that particular user um so how do we do that so we do that again looking at all the features that we have from the stories that is questions and answers and we feed that in into learning to rank algorithm um so future engineering I've mentioned that already in the answer ranking example it's a key issue to most of our models and I think it's a key issue to machine learning period And I know there's some discussion uh back and forth with for example deep learning practitioners of whether future engineering is alive or is dead and um it was funny that I went to Montreal uh a month ago and I had a presentation just after Joshua benjer was talking about all the Deep learning approaches and I was there on stage talking about how important feature engineering is and I think it is important and uh if you're interested more in this topic there's an answer I have on core about will deep learning replace all the other models uh where I say no and I refer to feature engineering as one of the reasons it won't um but anyway feature engineering is very important in the case of feed ranking we have a different number of features that relate to the user uh the story The interactions between the user and the story and all of those goes into the personalized learning to rank approach now what kind of models can you use to uh train a personalized learning through rank approach this is also a a long discussion but um the idea is start as simple as you can uh simple linear classifier like logistic regression might be good enough for you and you can build a personalized learning to rank approach using uh just a pointwise classifier like logistic regression um then you can go into more complex models so to speak and you can learn known linear interactions between your features and decision trees are really good for that and obviously that takes you to the next level which is tree ensembles like random forest and gr and boost in decision trees and you can go beyond that and then go to like what is proper learning to rank approaches like for example lambdamart which is an extension of gr of boed decision trees but includes you can optimize the loss function to a ranking metric um so we've we've tried all of those and basically we just go with whatever works best the good thing about having a nice infrastructure is that you really are not married to a particular model you basically can try all of them say okay what's the simplest one that I can use that gives me the best uh results and the simplest one is important because you want to be able to innovate and to grow and experiment on that model as much as possible but there are sometimes that you need to go to a more complex one there's no way around it because it's going to work better simply um and we do have C++ training Co U code that we've implemented ourselves but we also based a lot of the work that we do on thirdparty librar especially on the prototyping sort of like experimentation uh face okay and my last example of relevance it's email digest um I just included this one for this presentation because I every time that I've talked about things we do at quora there's somebody after the talk said hey for me the most impressive thing is when you send me the email and you haven't talked about this and say okay I'm going to talk about email diges so email diges the truth is um it's very similar to the feat uh example U although it has some differences right but the goal is still to present the most interesting and in this case answers and we limit the subset to 10 and also second uh goal is to optimize the number of emails and the times those emails are sent to each person so it's another personalized learning to rank approach but the difference here is we only rank answers no questions and we can do offline ranking as opposed to what we were doing in feed which needs to be real time we can do offline and we need to worry about saturation in the channel meaning like we can't send you too many emails and we need to understand when are you most likely to be interested in that email versus like just sending you tons of emails okay and oh ask to enter that's another example of relevance which I think it's it's pretty unique to uh Kora there's a blog post that we wrote about this um couple months ago so you can see more details there um a2a or ask to answer is this idea of we want to make sure that whenever we have a question we the system understands who are the people that we have on Kora that are most likely to write a good answer to that question and then once we know that we can suggest whoever wrote the question like hey this is a very good question you should ask that person over there because we know that that person can give you a good answer uh there's this there's the a TOA the manual approach the system itself also sends those requests on its own um so the idea is given a question and a viewer we want to rank all the other users on how well suited they are this is a kind of like complicated uh problem formulation because you have content text users sort of like kind of like mixed together um well suited is also complex because it has has a combination of the likelihood of whoever is viewing that rank to send the request plus the likelihood that if the receiver gets that request they will provide a good answer right so all of that goes into the formulation of the problem um it's really an extension of CTR prediction because you're you care both about the probability of the uh viewer sending the request but also the user responding to it it's kind of simp similar to the um there's there's some literature on sort of like these bidirectional recommender systems uh a typical example of this is the all the dating sites where you have to recommend a connection that you feel like the other end once it's made will also respond positively so this idea of connecting people where both ends need to actually respond positively um there's some research on that uh okay so I think I'm going to skip this because I'm going to run out of time and I'm going to move directly to uh other examples we also do recommendations on different things that help us sort of build a better graph right when I talked about the different overlays that we have and different data uh relations that we have I talked about users connecting to users users connecting to topics and so on so forth those things are really important uh but we need to actually push our users and we need to push the community in order to grow those relations so one way we do that is by recommending those connections right so if you're a user especially if you're a new user we want to make sure that you connect to topics so we'll give you recommendations of other topics that you should follow and those recommendations are going to be based on anything else it's going to be based on other topics you follow on users you follow obviously if you start following a lot of users that are interested in machine learning we'll say hey do you want to follow the topic machine learning um user interactions and topic related features in a similar way we'll recommend you users and we'll use anything else in order to recommend you users including and this is the the inverse the topics that you follow right if you're following machine learning and data science and related topics will look for the good users that relate to that uh area and we'll recommend you to follow those users and finally trending topics so this relates to this notion of personalized timeliness there's topics that come and go and appear all of a sudden and we feel like you might be interested on and in this case this uh there's this notion of we want to recommend a topic but this topic did not exist an hour ago right so it has its own challenges but we basically apply this notion of there's going to be different trendiness coefficients the global trend is like how much that Trends overall the social trend is how much that Trends in your social group and your user interest and so on so forth Okay so we've talked about understanding uh quality understanding relevance and making all sorts of personalized recommendations now we get to understanding relations um in the text and in uh all the content that we have there's a lot of relations that we need to understand like somebody was asking me about similar duplicate questions I'll talk about that in a second one of them is we want to understand what's the relation between a question and a topic and that's what we call topic labeling um this is a very interesting and I would say fascinating problem in our case because we have like a huge and we think very good topic ontology that goes all the way from like very broad topics like science to like super specific topics like tennis courts in Mountain View right and all of those things are available and once a question comes in we need to have an automatic machine learned algorithm that says hey given this question these are the topics that should go with that question um It's A Hard problem because we question by definition is going to have like very few words so we need to sort of like apply the all the knowledge that we have from like the few words that are available in that question to infer what are going to be the good topics uh that's it right now I mean I think the accuracy is pretty good but we still rely on users also entering topics and correcting sometimes the topic labeler because the topic labor made a mistake so the user will get notification hey these are the topics we think are the good ones you want want to add or remove some of them related questions so given interest in one question what is the next thing we should recommend you um and this is kind of like a combination of structure or um relation in the content but also a form of relevance so to speak um and the reason for that is because whenever you you hear any algorithm that talks about finding similars or finding related things uh people usually usually the naive approach is like okay let's base this algorithm purely on the content right it's like I find similar songs it's like okay what's the beat what's the harmony who sings that song let's find something similar the truth is in Practical applications the notion of similarity or relatedness and I'm speaking from my experience here but also from Netflix and other uh uh friends that I have many many companies is that it kind of like um it's confounded with the notion of interestingness as a matter of fact if you recommend two things that are extremely similar chances are the person is not going to be interested and like yeah okay you're recommending me the same question again I already know the answer to that question I want something that is related but it's also interesting uh and you can you can optimize your algorithm to that and you need to take though things into account like co- visits who visited this and also visited that which is something different from the purely textual or topic similarity um duplicate questions it's the extreme case of related questions and actually we don't like duplicate questions we love related questions we don't like duplicate questions so that's an interesting thing why don't we like duplicate questions because duplicate questions are basically diver diverting the energy in the system if you have the same people answer answering over and over again the same question or saying well this is almost exactly the same thing as I answered before or you have 10 answers on this question 10 answers on these other question which basically are the same um you don't want that to happen so um in this case the solution is we have a binary classifier that is trained with labeled data and it's similar to the notion of related questions but in this case it's purely attempting to to use label data on things that are duplicates and we know that they're duplicates how do we know that they're duplicates well because we've been through the process before we actually have a um a feature in our site which is called merging questions users and moderators and uh ourselves can actually merge questions when we decide that they're the same so that's a good training set to understand like what is what is things that were duplicat in the past and we use that to train using things like vector space models and usage based features okay user trust is another form of relation that we have in our Network and that actually this inputs into many of the other algorithms that we have we want to understand given a user how much can we trust that user in a given topic so we have this notion of topical trust and or trustworthiness um we take into account how many answers the user had written on the topics but also we want to understand how what was the quality of those answers so we want to take into account how many upbs and down boots those answers receive endorsements uh it's another feature we have in the site users receive endorsements on topics so this is kind of like um page rank approach where you have a network of connections and the trust propagates but propagates in a topical sense so if you have experts on machine learning endorsing or up boarding other exper machine learning that's going to propagate through the network okay so that's that was all about uh different kinds of applications on how to extract Knowledge from text and the other relations that we have in our data set um what are the models that we use um I don't have obviously time to go into all of those but I've mentioned several in different context but this is a I would say incomplete list list but there's like most of the things that we use currently in production that go all the way from simple logistic regressions or elastic Nets to more complex deep neural networks uh Matrix rization and other collaborative filtering approaches all of this fit into different things that I uh just explained okay so with that uh my conclusions is um quora not only has a lot of text but we also have complex data relations that actually help us understand more what's in that text and how much quality there is and how it is related to other things and so on and we need to have algorithms that understand that complexity and are able to optimize the dimensions that I talked about quality relevance and demand and machine learning is definitely one of the keys to our success and we have a big part of our technical engineering and data science team working on machine learning related approaches as you can uh see from what I've just presented um but we also have many problems that are still unsolved right if you think about the things that I mentioned we're still leaving a lot of things out that we could definitely do a lot more using machine learning NLP in order to extract even more knowledge and more relation from that text so I don't know if there's time for questions if they're not I'll be here here during lunch and you can also ask me on Kora because I'm like constantly answering answering questions on Kora thank [Applause] you