Devreal

Scale By The Bay 2019: Omar Alonso, Fast and scalable domain-specific knowledge graphs generation

Scale By The Bay 2019: Omar Alonso, Fast and scalable domain-specific knowledge graphs generation

Recording: Scale By The Bay 2019: Omar Alonso, Fast and scalable domain-specific knowledge graphs generation

[Music] hello from Microsoft anymore I used to work two months ago now I joined karai a startup so this talk is about a lot of things that I used to do at Microsoft but at the same time and we'll be talking about what we're working on at this new cool place all right so the outline is to give an overview why knowledge graphs are important nowadays and then I'm going to go through three examples of Nora's graphs that I'm working that I work on and I'm working on the first one is one based on social data in particular Twitter the second is on e-commerce and the third one is on healthcare and then we'll wrap up with conclusions so notice graphs what's a knowledge graph is basically a graph where you can describe objects of interest and connections and we're interested mostly in organizing data as nodes and edges and if you use any search engine like Microsoft Bing or Google you're using the knowledge graph under the cover so every time you see a curry understand entity on the right you know the name of a person the age what they do etc etc that's kind of power by either Satori which is the Microsoft knowledge graph or the Google knowledge graph if you happen to use Google this notion of building knowledge graph is taking off in many other companies a year ago Amazon announced that they're building their Amazon product graph and there's a big team in Seattle basically building a graph as well as avera doing the same on at the same time has been a lot of work in the past many many decades under the umbrella of knowledge bases so you may be familiar with iago which was done on max planck other things like freebase here in the Bay Area that then later was acquired by Google the terms knowledge graph and knowledge basis are mostly used interchangeably in this talk I'm not going to describe what the difference between them I'm going just going to say that the focus in this short talk is mostly on the contract construction of generation of the graphs and not so much on the reason second part is we're interested basically in scalability and fast generation which means can we scale the construction of all these graphs and how fast can we do this so the goal is to identify nodes and derive those edges and how can we do this automatically so far most of the research of knowledge graph and knowledge bases they usually use Wikipedia as a source so this is very good because you you know the benefits of Wikipedia the content is easy to read is well-structured the wiki is easy to parse there's Wikipedians we're going to edit and so forth the drawbacks is obviously coverage not every topic is in Wikipedia only the head topics or the important topics on Wikipedia a lot of stuff is outdated and if you get into real specific things like biases there's a lot of bias in some of these we'll give me the articles now the problem is what do we do when there's no Wikipedia so what do we do if there's no Wikipedia and there's plenty of examples in real life what you don't have Wikipedia and if you don't have Wikipedia how can you derive a graph and does this talk the first one is SKG the social knowledge graph we have if you're really interested in the topic i have pointers to papers that we published we have one in CIA and the other one on JC the other kind of get into the details of all of this but high level overview observations on working with twitter data at Microsoft we used to have access to trader for about 10 years 10 10 years which will be paid for the firehouse and we store all to the data and we do a lot of stuff on top of that so after working for nearly a decade on pure data my observations are the following so drinking from a firehose is very very hard and very difficult there's very low level the server does value of the atomic level on a tweed and a Facebook post on Foursquare check-in etc etc but at the same time there's a huge potential to build new artifacts and this applies to aggregations going beyond recommending tags or friends or accounts or trending topics and the idea was how to build is possible to automatically generate a knowledge graph order to the data and I firmly believe that all these ideas are applicable to other social networks like LinkedIn or Facebook all right so a squeegee the input is very simple the fire hose to the fire hose and the output is a knowledge graph simple right so we're gonna identify four components of our graph so the nodes are going to be the links the topics the entities in this case people organizations from places and then over time and I'm going to go back to this and tell you why time is super important in social networks so we're going to look at the graph from these from those views and we're going to focus on super high quality content relevant content from trusted users and good topics and I'll go through these as well this is a core schema looks complicated but I'm going to choose highlight the main things so we have users post topics and links and imagine this is kind of an OLAP review if you want to call it or if you're very familiar with Excel the imagine you just read an entire fire hose in Excel and then you play with pivot tables something like that right so you have tables of users everything that Twitter's going to give you plus things that we automatically derive like static rank Authority score then we have the links the same you know what's the static rank of the link the topics on the post this is important we because we call them supporting evidence or provenance as people were covered in databases when you want to check a particular edge or node then we're going to build a lot of different subsets one of them is the notion of verified users and trusted users so if you're in Twitter you have this check check mark that's your verified but it is a manual samanya process from Twitter so what we did is we build on our the activation network that will just take the verified users and automatically scale the number of users that we call them trusted so if Bill Gates who is a verified user add me in Twitter and I add him back then I'm a trusted user so then if I add Sophia Sophia at me back then she's a trusted user on ring number two so from Bill Gates a ring number zero I myself ring number one Sophia Rings to etc so in a few hours we have a massive set of trusted users then we look at links so we have good links and then we have viral links and then we have trending links and if we dive into topics a topic should be a hashtag could be an entity an entity here means using an identity extractor to tag to extract tags from from a tweet cache tags which is basically the dollar sign next to MSFT will mean that you're talking about Microsoft in the context of financial information versus the pound MSFT which is the topic Microsoft and then other things like engrams trending and a bunch of annotations all of these to say that you know we take a massive fire hose and then we shrink it into this kind of four nodes abuser post topics and links and then we build a lot of different connections which allows you to give it a use here we can see what are the links what are the topics examples of tweets given a link who are the users who are sharing this link post a retweet this link and so forth alright so far so good we had a graph what do we do with it so here's the application here's kind of a cool application once we have a graph we want to basically build an application that will tell you the evolution of a story so here's the pitch you know there's a lot of social post on something that they will make it into a few news articles and some of them will have a single one Wikipedia page all right those are thousands opposed that in ended up into a few hundred news articles into one Wikipedia page that's the planet a few editors - maybe one or two Wikipedia's you can guess obviously the bias and coverage and all those things but that's not the point the point is can we leverage this human scale his kind of human sensing on the planet and then combining with wiki and can we automatically build you a Wikipedia page based on social data that's the application so we're going to use Twitter as a distributed human crawler all of you guys are sharing links we're going to extract all the links from Twitter as part of a backbone of the story and then we're going to put a Wicca fication I got him on top of so once we have the graph we will run the Wikia fication I'll go on top of it and then we'll build you this thing which is kind of a database of stories so it's a synthetic document I'll give you another example later that gives you provenance supporting evidence evolution and different pivots to play with different language and you can archive otherwise you can kind of archive what was said and the the idea why we wanted to do this because we wanted to kind of get an idea of the evolution of the u.s. election so white Ram was selected as president everybody's been or happy or sad by wanted to see what how we came out Java to electron so that was the evolution of the story so here's the processing pipeline we're gonna have harvest the data we build a graph and then we're gonna build a couple of extensions to well-known information retrieval techniques one is pseudo relevance feedback so when I expand that also the query expansion so we get the graph then we have all the links and then we expand all these things with extra metadata and then we're going to build the about the story on top of it for that the the idea of counting votes or likes or retweet this kind of shallow doesn't give you a lot of signal for all these things so we built our kind of different voting destructor here where if you just use basic counts as I said before it's very artificial the popularity so so then we're going to count by different things we're going to count the links the currents of the links with the hashtag the currents of the hashtag with the link with the person so for us we got all this combination of different ways of computing votes and imagine like a record of different ways to count different votes once we have this the second part of the of the puzzle is to build a hashtag index so just an index of hashtags and an explanation for those hashtags we like this so we have the hashtag index and then if you're looking at hash recount 2016 if we're looking at December the second which is the first entry is mostly about some issues going in Wisconsin about the recount and then hashtags are the devote Wisconsin recount et cetera if you will go for one day the recount is to a different state in this case Michigan so those are the the timestamp the entry the contextual vector which is kind of a signature of what people were talking about that hash tag the signature which is similar idea kind of a summary of what the link was share that's the URL and then these related hashtags it allows you to pivot so you're looking at recount 2016 you can say oh there was an issue with the Wisconsin recount or the Michigan recount and you can go to that story so you cannot make in those linkage automatically same way we're going to use this to expand the notion of pseudo relevance feedback and query expansion I'm just going to give you this but it works as is for example Super Bowl is very popular hash tag but its equivalent of SB 50 was only on 2016 and SB 51 was only valid on 2017 however Super Bowl was valid for both year so the lesson learned here is in social networks time is crucial because context is only valid for certain amounts of time so one of capture is we're going to capture all these things for the different expansions we're going to use sim has which is a clustering technique for grouping similar hash tags and then obviously we used to have access to the Bing search / lock so it's a lot of ways on a search engine that give us you know training things and so forth anyway all these two sadism is a big linear combination to give us a score and allows you to do things like this if according to tradition the query Clinton in July 31st one of most popular hashtag was Hillary Clinton and that's the vector definition for that day but at the same time on that day she was mostly mentioned in context of Benghazi that's another definition for hashtag Benghazi on the 27 then there's Clinton Twitter's email scandal call me resignation that's the relationship if you go to Trump by the way I'm just using your selection because if you're really into data mining and all this crazy stuff please download the US election there's plenty for pretty much for everything you want to do us as always the day I said to go on Agata forth from train related to make America great again then on the 16 September is from nation from army political stuff obviously Bernie Bernie or bust the Millennials should have been Bernie and then still Sanders all these things you are kind of automatically discovering those relationships because people are just kind of tweeting or or linking things and under the graph were discovering this linkage automatically and in the final node moving away from politics on October 3rd obviously Star Wars that's rock one but because Carrie Fisher passed away then a rock one was associated with recipients princess all those things are detected automatically why I'm pitching it is to you guys because all these things is very hard to put in any Wikipedia documents or have someone in company in your company to do all this linkage for you so here which is kind of piggybacking on top of the user base once we have this the rest is building the story and the story is traversing the hashtag index looping through all these tables we ranking the links and create the story and this is kind of the summary of the story it said this is the story for the US elections looks like a Wikipedia page because it was derived using the same style as a Wikipedia page but under the covers is all data coming from Twitter on top of the graph so the table of content is heatless clustering then the story evolution generation is for the to summation to you then we have the related stories because we making these connections this could be the the hashtags related to other hashtags the sources because you're sending a link then we can detect the source that domain of that link and that will be the source of the reference in the page so we're not making this up this is the story told by the planet so it's not Washington Post view or the New York Times or somebody else's the planet is giving you the story which is good because they will have all these different perspectives which is I think very useful and because people are looking for this the story already the queries so the quarters already know how to fire and when to fire the story sanic this is a comparison of the cinema fires from our page to the official Wikipedia article and we're as good as the Wikipedia article and all this to say that a year ago you can fire this and Bing and what you see on the round and was basically the story that was popularly done powered by the graph alright that's graph number one graph number two again another publication we're gonna switch context and go from social media to e-commerce so here is I'm going to look at brands products and categories seniors are you issuing the query for a brand Microsoft and you would like to see all the products office surface windows or maybe you're looking for a product for example jeans and I want to get brands Calvin Klein Levi's or I'm quoting for the category smart phones and I want to return i phone and galaxy or pixel this is interesting because a product can be an item or for sale or a service it could be a leader samba or could be an insurance policy or it could be the food delivery if you do this you can detect competitors the fact that Nike and Adidas both both in your cells running shoes then I can derive that Adidas and Nike are competitors and in e-commerce scenario like for example serving ads this is very very important so then I can I can detect related products greater brands products amélie's and so forth it sounds simple and trivial but there's a lot of different challenges here the first one is this is a very dynamic domain so you got different brands and products they appear and disappear as we go there's a lack of major sources with clean brand and product data so wikipedia will cover Nike and Adidas but for the torso and tail of the distributions there's basically nothing absolutely nothing it's very hard to define on the tech products and sometimes the distinction between a brand and a product is kind of blurs very difficult to see and retailers don't always going to provide clean data so what's our approach is unsupervised Tzar going to focus on data quality and simplicity we're going to generate brands using data fusion schemes they will attack the brands with categories and from that we derive the products so a nutshell works like this we use in a voting mechanism so given n number of different sources we're going to each source will vote if they see this particular term or token and if this goes above a specific threshold we can derive with some confidence that adidas is a brand or Nike a zebra once we do this we go into the search query logs and try to extract the most frequent domain at position number one this is good because give us for example for a particular term apple an alias for example Apple Inc Apple Computer they're all pointing to Apple calm which seems to be the dominant domain for this brand and it's here a score of 1 means that all of the sources understand Apple but if we go down the list you'll see things that people are not very clear particular for small merchants which will have a presence on the web but they're not super popular once we have the brands and the domains the second step is to derive a category this is also very very hard so how we categorize Apple do you think Apple is smartphone sees electronics it's also fashion you know why because they sell watches right that's hard so we start from a brand then we extract all the URLs when a search for queries the point to this URL from those URLs we're going to extract all the categories we're going to do an aggregation on the categories at the end we're going to derive a list of categories that are associated for a brand we're going to compute the lot of different probabilities in particular the probability of a category given a brand the probability of a query given a URL and making again other user formula we're going to get what's the list what's the list of different categories for a brand and with this we can also give you the popularity of a brand once we have the brands to the carriers and the maze the last piece of the cake is basically how to derive a problem so with the right products from two ways one is from retail catalog so say walmart.com kind of dumps data continue and then we have to basically identify which are this Ingram do actually match a particular brand and then use other things like for example a deep structure semantic modeling which is Microsoft thing on top of the product ads to detect some of those names then we use k-means clustering with a bunch of heuristics because the larger clusters are very large and then at the end we generate a set of M different levels this is just using the catalogs just another set of techniques for using bitter keywords the keywords that an advertiser will use to bid on particular search query terms that's another different techniques Ingram's language models based on KL etc etc this also tell you that from nothing from different sources we can kind of one by one below these little components and at the end we can give you a graph in terms of the system design because we're using a wide range of data sources you can now the new there are sources oswego and they're not going to affect the voting schemes and how we derive all these things this is a very bottom-up approach with a focus on simplicity and data cleaning so we have to have data cleaning as part of the data in the graphs who are trying to convert out everything is not so clean this same as the previous project with social data on Twitter these are all implemented in cosmos and scope which is the Microsoft equivalent to MapReduce in Google they work very well at scale all right third graph to be knowledge graphs for healthcare so I talked about social data Twitter you're mostly folks here familiar as well as you know ecommerce Amazon but here this is a different beast why because there's very strong scientific medical knowledge out there there's a lot of existing taxonomy and there are sources available for this for examples you have sno-med which is the systematized nomenclature for medicine you can download there is Erik norm which is basically a taxonomy of all the available medications in the US then you have mesh medical subject heading sense and point to other different sources and taxonomies everything kind of looking at a particular view of healthcare the challenges here are the following one is this is extremely sensitive data so it's not a tweet okay it's not a link I don't care about hashtag mega this is your medical records right so this is very very sensitive and if if you have access to a hospital you have medical records will be great and will be fantastic to kind of drive kind of a knowledge from these things the third item is the notion of relevance in healthcare and we call this clinical relevance so you know relevance in web you know popularity relevance you know Instagram or Twitter but what's relevance in a medical domain so we're going to the emergency room for something you know the doctor will check recheck you to give you context what happened check your records to ancient and then make a call on what things are relevant or not then there's a vocabulary mismatch problem so the doctors they're very precise and the terminology but we as patients were completely imprecise you know hurts here I don't feel well etc so you have a mismatch in terms of terminology and then if you want to do machine learning on top of all these things you know dinner labeling and curation in this particular domain is super hard not only because of the sensitivity of the data because who is going to label these things we're going to hire more doctors who is you know label so can we do your active learning label procreation and all those things are super hand in in these domains and what are we trying to build a Qi I just started a couple of months ago so I don't have a lot of beef to show you at the moment although Sophia is going to give you a great presentation on human loop but we're trying to build a comprehensive medical knowledge graph that will cover diagnosis symptoms descriptions etc as well as labs and medications if we do this and we're doing this we can run question-answering on top of it and we can produce automatic diagnosis which is what to them part of of this scheme is to have a very enriched human-in-the-loop data pipeline that will have the knowledge graph under the covers a bunch of email and algorithms but also kind of a health coach or MD medical doctor in the loop to hopefully give you the best diagnosis possible to conclude and I really rush this presentation so happy to take questions or talk offline this is a super active area in academia industry if you're subscribed to the CI cm there's a great article in August this year where they talk about the importance of knowledge graph and the article is co-written by Microsoft Google Facebook and eBay just to give you an idea how important this area is there's a lot of many workshops on knowledge bases and knowledge graphs in top-tier conferences and in the field this requires a combination of many many different techniques so large scale processing formation retrieval machine learning ai NLP 11 etc so you cannot do yourself by yourself you have to work in a team who is very diverse and at the same time which is super important you have to have the main expertise because you're not going to build a KB or a knowledge graph of the shelf and then sell it so you have to do it in context of the domain and sometimes either you are the expert on the domain all you were working with are my experts on this and that's about it thank you so much for attention [Applause] there's time yes any questions thanks for the talk had a question about your knowledge graphs for generating for random product you talked about generating representative labels for products like peering down on putting down on a larger set of tags or words that you consider should be represented at label oh where some of your concerns and trade-offs when you're choosing something that could be displayable or representative you're gonna have you're gonna have a lot of different labels and sometimes some of these things are unclear in particular when you have models and versions in case of apparel for example you know this teacher would be in a size M or L sometimes the jeans you know shirt is good as a product but sometimes you get extra information that maybe is a bit redundant or doesn't help you a lot because you may detect many many products that the variations are on these little things particular quantities or models or variations on the other hand sometimes you have super high labels like in your shoes and everything you know could be shoes can you get into more detail so that's that's something that the team is still working on but when we started working on this that was one of the other things that we notice so if you if you do the top 5 for example of the products for Microsoft they were very very good but if you go down the list some of the things are like oops near dupes and a lot of different versioning or extra things that don't add a lot because a company won't sell like you know millions of products usually is this kind of a small set compared to everything it's para question some the knowledge graph you're trying to build it seems to be from like the medical textbooks and the EMR information we kind of give two distinct sets of information looking to try to merge them into one for like a kind of a holistic picture of kind of health or ways to help individuals or say via so they did that we're working on is to build kind of a unified knowledge graph of health care like I said I just don't claim it that we have everything ready but we're it that's our direction so there's a graph there's medications and drugs that we're trying to ingest once we have these drugs so this graph of diagnosis and symptoms and also questions which I forgot to to mention there's many ways of asking if you're sick or not there's many ways or telling your doctor you're sick once we do that imagine you have this graph and obviously you can kind of ingest part of the graph into an EMR or make it an EMR an application on top of the graph so we have first beam which is our application though you know it's available if you want to play with it that basically uses the graph under the covers and has launched in a new version that will have some of these autocompletes that are like more fancy but that's the way that's a plan of cure that's what Mexico happy to take questions offline thank you very much for your attention in summer [Music]