Devreal

Aggregations and Knowledge Extraction from Social Data

Event: Aggregations and knowledge extraction from social data: challenges and lessons

SF Scala: Omar Alonso, Aggregations and knowledge extraction from social data

Recording: SF Scala: Omar Alonso, Aggregations and knowledge extraction from social data

[Music] thanks for coming on this rainy day rainy night and the talk is about how to build aggregations and knowledge extraction from data and this is the typical disclosure that is just basically my perspective and not a policy official policy or position of Microsoft all this work is joint work with current and former Microsoft colleagues and also interns so I'm going to give a brief introduction to the problem of reading from the firehose meaning reading and learning from data sources in particular social and then I'm going to talk about three projects how to search for things to do in any city on the planet searching the social web and how you can generate a knowledge graph a lightweight knowledge graph out of tutor data and then I'll draw some conclusions everything by the way a lot of the stuff hasn't been published so I'll do have links to our papers if you want to read more all the papers have all the you know super technical details on how these things were done so what's the landscape in terms of social media you're all familiar with Twitter which is basically the place to go for micro blogging Facebook is social networks photo sharing is Instagram Foursquare is an example of a location-based social network and kora is an example of search and answering question answering all those all these different platforms or their sources share the same commonalities which basically they all have a social graph mostly followers followers they all have recommendations they recommend you friends accounts topics links point of interest they all have trending so there's training links during the hashtag trending topics etc however drinking from a firehose deck is really hard right and there is a very low value at the atomic level meaning like there's very low value the tweet a facebook posts a single Pio eye etc etc now because there's a lot of human sensing on the planet at any point in time there's a huge potential to build new artifacts in particular aggregations and this talk is about how we build three different types of our Galatians the first one is searching for things to do so we're going to connect to crowds so you are in San Francisco and you're checking in at the Golden Gate or you're checking in at MoMA while you're checking in you know at this place and then at the same time there's a lot of people who are searching for things to do in San Francisco this is typical you go to the autocomplete in Google or in Microsoft Bing and if you say things to do in it's a city okay this is like an example of that all right but the project in the net suit is basically how can we use location-based social networks like Foursquare for example that data said to build a recommender system so how to how to use POS checking's to suggest Bo eyes of place of interest to go in any part of the planet so but it's a check-in in case you don't know if you're using Foursquare of overuse in Facebook for doing checking's you check in a particular point on the planet you just press a button and then the platform will record that you were here at this particular lat/long and at this time these are two examples of you know different ways of doing check-ins and what you can do with that the beauty of a check-in is there's a temporal view and because of these you can see different human patterns so for example we know that Disneyland it's popular 365 days a year but if you want to go to your semi tee you have to go on the summer so the chickens will give you all this I will give you human mobility correlations with Cesar's different parts of the planets etc now whatever if we do that to be the recommender system so we're not target what we call a recreational quarries so recreational queries is a common query intent so the user is looking to perform an activity anchor around a place or a city so things to do in San Francisco things to do with kids in Seattle romantic things to do in Paris and this is to be for both tourists and local users desktop and mobile are their scenarios so how can we extract location information well I work for Microsoft and therefore we have access to lots of data one of them is the query logs so everything that you're searching on bing you know gets told in a query log so we can go to the query logs and get an idea of what are the queries that have different patterns like things to do places to go where to etcetera etc so we can see a sampler who's going to sample different queries with different patterns and from that we can get all the queries that we believe contain a location so that's the second part of the of the chart and then we run an immediate extraction of locations so then we are going to annotate all the quarries with location information and based on that we're going to identify all the queries that contain the patterns and location data we can do the same for other things like questions if we're coming from Chora URLs from the IE browser logs etc etc once we have that then we have a massive data set of recreational queries that then we use to build a taxonomy and that is data driven so let me dive a little bit more on that so first of all we get core logs as I already mentioned all the kulaks will tell us a lot about you know popularity of cities and so forth we have behavioral data from the browser so we know if people using for example ie 11 or 10 visited a place that is a candidate for to go and then we have the Foursquare check-ins Plaza tips so the tips are if you leave like a tip or review on a particular PII we can get that as well and then the recreational query taxonomy although this is kind of a very shorter screenshot you have a query and then you have a query geographical constraint temporal constrain and activity preference suitability and other constraints in the activity for example you can say things to do in San Francisco which is you know very specific but also you can say things within the Bay Area which is more high-level there's absolute time so you can say things you're in San Francisco tonight things to do in San Francisco tomorrow the preferences are things to do in San Francisco too naive the word to go were to go shopping if there's an age group if you're looking for stuff to do with kids with teenagers if you're looking for if you're concerned with budget you can say cheap things to do in San Francisco anyway all those things were done mining the koi logs I'm building these taxonomy we call them aspects so that the other part is once we have the quarry classification and the aspects then we can classify a map particular queries and point of interest of supporting evidence for this entry on the on the taxonomy for modeling there Pui relevance we use in a maximum likelihood estimation and they were used in poorer ranking which is the probability of check-ins during weeks and seasons so there are certain places are more popular more popular during the week other procedures are more popular during the seasons a specific seasons as well as week we can't and then for each of these aspects we can rank them so we can filter all this aspect for example like eating drinking romantic kids etc and with the tips we can match this particular tip for example this is a good place for kids with the entry in the taxonomy that is a constraint for kids I know this sounds a little bit crazy but you know I'll show a few examples now there's a problem with a ranking Pio eyes or attractions and Pio eyes the first one is not everything is good as a recreational career for example a lot of people will for example do a check-in and a bus station or the high school office gas stations prisons etc not not very good we don't want to use these guys at the same time there's also things like a bridge for example a Golden Gate is a very good place in San Francisco to visit which has a lot of check-ins but also other bridges in particular if you want to if you have to cross the bridge because you're commuting people not be making a check-in because there are the toll plaza or something else as a signal to say I'm late or I'm getting there it's not very useful to use in the same way train station there are certain registrations super popular otherwise just a train station potentially good examples could be restaurant churches and parks and obviously good examples are famous landmarks and tourist attractions this little example I mean to identify those four categories took us quite a bit of time to understand the utility of Pio I so alexei was mentioning the notion of labeling the utility of a point of interest is not straightforward once we do this we collect labels with our UHS which is our internal Mechanical Turk within Microsoft with three ratings if the the recommendation was good for bad we use fats rank another internal tool standard TLC ranking gauge TLC is the Microsoft equivalent to wicker and a bunch of other features for example core impressions categories and other things like for example repeated visits versus new visits and it looks like this if you go to Bing and you say at least when I took this picture things to do in San Francisco this is the recommendation ok straightforward things to do in San Francisco with kids in the second separate recommendation you got a question so your choices are on Amtrak I cannot comment on the recruiting of people however it's a tool that also allows anybody within Microsoft to crowdsource so it's not just internal eaters anybody can get labels even within your team you can use your choice to gather your own labels it says yes yes all of that and a bit more yes not now not in this one oh yeah the question is what are the details of the internal Microsoft crowd sourcing tool I mentioned we have you hrs it which is a comment here that we got our labels using you hrs which is the internal improvement and the question from the audience we'll see if I can share some lights and UHS and I say that you know I cannot share light on the recruiting of of people who work on your choice but I say that you can crowdsource you can work and you can also hide editors that's all I can say all right so this is things to do in San Francisco thanks to the San Francisco week it as you can see it's a different recommendation now we can go live it up we're going to say things to do in California okay landmarks from California thinks you're in the United States one level up and once we can do that so this is like this is what how it's being used in the Microsoft Bing search engine when this works for other cities you can say things to do in Barcelona things in Barcelona with kids I'm not gonna them on that but here's something different we call urban maps this is an internal tool so it's not in production yet so this is an example of Seattle if I'm looking for things to do in Seattle an evening that's the map of Seattle and then it's going to highlight like a heat map were the activity on the city is and then it's going to show you different places to go so for example if you happen to be at the Pike Place Market you can go to the crab pod you can go to Starbucks you can go to downtown Seattle or you can go to the Hard Rock Cafe they're going to give you in tenaris or were to go if you happen to for example do the same query things to do in Seattle at night or for example places to go clubbing the the map is going to change I'm going to show you a different recommendations and your example another set of recommendations so the point here is here's an hour an aggregation on how you can derive insights using location-based social networks build a a recommender system and B how can you build a different experience you've seen maps and we call these urban maps so if you're if you're in a in any city in the planet and you can see where the action is depending on the day or depending on what you're looking for you can you can do this all right so that's aggregation number one aggregation number two is searching the social web this is basically in a nutshell how can you search within Twitter or within Facebook the point is not for searching post but the punished for searching links that are being highly shared within a social network so there's a lot of there's a mini web within every social network a lot of people are calling this the like economy why because there's a lot of people liking or sharing or reorder or sharing URLs whatever if we only index that forget about the rest of the web only that alright so we're going to build what we call a different type of search engine as I mentioned the idea is to take advantage of the social web structure also call as the like economy of social or social buttons if we go to a page you usually will see a button that is like Facebook or Pinterest or Twitter or LinkedIn what have you there's always some social button this particular project is links that have a lot of chatter on Twitter we're going to extract users hash tags and keywords I'm going to rank all the links based on a combination of virality and popularity and at the end we're going to build a Wayback Machine so how can I search in the past so what people said at a particular event so this is a search in the social web so you issue a query for example hashtag climate change and instead of showing tweets we're going to show what we call social cards so the card the focus here is on the document as you can see here's a document coming from LA Times there's a picture there's a caption and has been shared by the 129 people and what we care is basically not only the accounts what people have said on those links all right well I was just trying to say that here we want the the content the people and then we also want related hashtags so we shoulda query for hashtag climate change but a related topic is climate excellent and climate change research this is an example the summer I the social card is basically the link information on top trusted users I'll explain what a trusted user is who share the link all the pivots allows us to cannot traverse the story all the post link annotations if you want to explore the user which I'm not going to show it today but we call that an operation called x-ray if you want to explore this particular account we can dive into the account and it's not just the handle of Twitter but everything that is - the handle has been done and then related hashtags as queries so the background the back in pipeline contains five main stages the content selection the user selection the link selection and what we call the final cut and then we do annotation and indexing and presentation so the first step which is content selection is a machine learning classifier that will select if the Twitter's interesting or not or is a good content this seems like very straightforward but actually took us a while to understand this so picture the following I'll give you 140 characters and you have to tell me if the studious has good contour not hard problem synth seems very straightforward it's a hard problem because it was very hard to label and understand how to label tweets anyway we have a publication on this if you're interested but we can stamp a tweet with a score indicating that how good is particular to it is we do a spam filtering we normalize all the links then we run mini - Fournier duplicate detection on links J for example if you tweet an emoticon low-quality if you say I'm attending this presentation in San Francisco I'd say dominant labs URL you use in 140 characters well written uppercase lowercase etc that'll be a good thing if you use a hundred of forty characters when we do this was before 280 everything is a hashtag then not useful everything is a single character is not user correct so if you if you if I pick you and I give you a tweet you're gonna you say this is in this is useful versus it's just junk and the point is you want to index only high quality content no low-level content and then we extract metadata the Twitter cars Open Graph metadata ODP link classification and a bunch of our stuff a link normalization means that for example you have because it's a lot of links sharing within the social network all these links will have different parameters so this is coming from a particular site so we have a lot of forms that are embedded in the links because people do not track where they share the link so we want to normalize to exactly the domain and the page that's all now a very important part is how to do user selection you're all familiar with this notion of a verified user which is Twitter's method for heinki reading people now this is not going to scale because the number of verified users is kind of low so what we did is we came up with an algorithm that is going to automatically scale up this process to what we call trusted users and works late the following works as follows so say that Bill Gates Bill Gates is a verified user Bill Gates adds me in Twitter I'm not verified user if I replay if I reply back to Bill Gates then I'm trusted in ring number one so we start with verified users in ring zero and then now I'm a trusted user a ring one if I add Alexi who is also not verified user and Alexia replies back to me then Alexa is now in ring number two so we can this basically an activation network and we can scale to what we did is with 17 million users in a few hours so from a very small set of threads and users we have verified users we have 17 million trusted users that we know at least you are good yeah there's more readers on our paper the second the second step is a Thursday base link selection we have good content we have good trusted good users now we want to get good links so we're going to look at all the tweets in a given period of time that contains links for example say that it's like doing it grabs we're going to grab all the tweets that contains links and then we're going to construct a diffusion tree for a link or identify sorry popular viral links we're gonna run some similar link clustering I'm going to do more filtering so when a filter on English only we're going to remove all the links that don't have any metadata and we're going to discard all the links that only match a domain this is interesting because in lot on social networks is a lot of you know ads sharing a lot of people promoting web sites we don't want that we want a domain and a page because we want that stuff and finally we're going to extract what we called contextual vectors this you may find this interesting so contextual vector is a vector of engrams from a set of tweets related to a hashtag or an entity this is not a summary it's not a summer it's kind of a lightweight signature of a link and the beauty of this is there's a temporal context and you'll see later on the talk how we use this many different ways so for example let's take a look at the contextual vectors the got actual vectors for generate a 7th 2015 that's the Paris attack so there were two trending hashtags Charlie Hebdo and justice Charlie I'm not French so sorry five mich means mispronounce this but for Charlie Hebdo we know that the engrams are basically free speech Charlie Hebdo attack satirical magazine said de Paris attacks for this visually it's kind of similar you know Charlie Hebdo free speech Trafalgar Square Saturday tourists attacked Paris attack if we look at January 26 2016 rogue one Carrie Fisher Leia actors princess you know Darth Vader Star Wars pre pre known but at the same time you know rest in peace princess you know Carrie Fisher Leia Prince's rest in peace Star Wars more all these things in context so the same idea social signature so contextual vectors given a hashtag or entity will give you an Engram here is given a link st. Rick ok so here's a link from Fox News for example and that's the the official that's the correct URL that's the title and that's the signature it's not exactly the same so remember that the signatures and the vectors are a summary of what the people are saying when they're talking about this link okay so if I go back to the previous social cars in case I wasn't very clear if we take for example those three posts there are comments on the links those that's the data that we use to build the context the contextual vectors and the signatures okay so it's not on the LA Times article is on what people are saying on the LA Times article all right so you saw Charlie Hebdo and princess naila now let's take a look at the first one from Fox News the second from Politico by the way we don't care if you are from the left on the right which is don't care that's another article on the recount that's the title and that's the an example of the of the signature and here's an example that should convey the concept can you identify this for me pre-clear so this is the video that was shared on the asana airplane crash that day it was only a video now the activity on Twitter that they basically summarized that it was crash landing the plane crash the triple 7 someone you know loved you were saying that's a pilot error we don't know it's just on a day it's just on this day this is what people said and that is basically a social signature so when I use all this to make an index on those things and the and by the way the beauty of these things for example is if you have images or videos which don't have any textual content the social exchanges are basically the documents for those images or videos now the results if we run this every day for 20 hours we will run this process we go you know we're gonna stream a lot of the content so we start with for example 227 minutes of tweet the user selection 19 millions that's from just selecting tweet then we go to 44 million links 18 million users 15k links 1k links after the final cut 25 million users 17 million after user selection 9k hashtags and 3k engrams all I'm trying to say is after we apply all these different mechanism you have a very rich data set to build something on top of it so instead of dealing with the entire thing all the stuff that you don't care which is low quality here's a much smaller thing that you can do a lot of cool stuff I'll show you a lot of these examples one thing you can do is you can build a Wayback Machine so I'm pretty sure you're familiar with the Internet Archive where you say hey how was my page five years ago we're gonna do the same thing so going to archive daily snapshot we're gonna take a look at the top 50 hashtags for 2016 and we have the following scenarios what was said on a specific day what was relevant on a specific day and we're going to harvest over time you are selections if you really want to play with social data I strongly recommend please use the u.s. selection because it's beautiful whatever you're looking for there's going to be an entry there so I'm showing you two what would say the day before the elections the day after the elections okay there's no political party here it is top hashtags for November the 8th top hashtags for November 9 if you look at the the top is just an example of for social cards mostly people are talking about a I'm going to vote maybe my candidate ways I'm going to vote if you look at below is obviously from one then the media say oh why we go wrong and then a lot of people say I'm leaving the u.s. typical right so the third one is all the actors say I'm gone you know I'm gone and it's interesting because if you look at I'm gonna highlight here Kyle exit alright so all of you were saying let's just leave California there was a brac said we should do a calyx it interesting so these are the things that you know maybe you know the Tory guys should be doing or the Facebook guy should be doing but you can go back and see what was say that day on a particular day how we do this against Wikipedia page so we're going to compute precision and recall against the official Wikipedia page of the US presidential election so go for some reason it's not showing up but anyway you can you can see on the slides later the official URL for the Wikipedia u.s. elections that's exactly day by day what happened we do the same so we're going to take a particular date then we're gonna find our hashtags for that day and then we're going to look up the Wikipedia entry and they want to measure if these are the same or not so if we look at the recall at 25 were 82% if we do recall at 25 with filter data is 92 percent precision at 25 86 precision of 50 gets better 92 percent and if we do recall at 50 is 93 percent this is completely unsupervised so we build this thing although Wikipedia is a literally driven you can also argue that our stuff is human base because it is written by people so the cool thing is like they're pretty pretty close the last part of the talk and I'm gonna give plenty of time for questions is how to use all the little components that I just mentioned and derive how a story evolved on how to build knowledge graphs so here's the pitch you can think of there's an event happening there's a lot there's a social post a lot of social posts there's going to be a news article and then a Wikipedia page so Wikipedia page in my opinion is where the event goes to die okay so didn't happen that's it there's a Wikipedia page shrink-wrap were gone what's new now social posts are thousands news are written by editors hundreds and then you have one Wikipedia page which means you have the planet you have editors and then you have Wikipedians so the world millions editors hundreds and then one or two wikipedia as per wiki page so can we do better can we leverage this human sensing at scale and at the same time combining social and wiki so here's this is our project our approach is basically to use social network as a distributed human computation crawler I'm gonna set it again every time you're on Facebook sharing link every time you're in Twitter journaling or LinkedIn or Pinterest or whatever you're basically crawling the web for free okay so I mean we had Microsoft there's a team who does just crawling basically you know traversing the graph to find pages in a social network people are saying this is good this is good this is good so let's just leverage let's just piggyback on top of that so we're going to extract all the links sharing Twitter that serve us a backbone of the story and we're going to automatically build a document once we build a document we have related stories and then because people know how to read a Wikipedia page we're gonna build a Wikipedia page for it all right it's like this it's a database of stories it's a synthetic document so you see on the on the right is basically interesting because Alexa was mentioned in fake news and Russia this is exactly that so there's an entry here's a story of someone about obviously about the u.s. elections and Russia intervention that's on the right someone wrote an article there's a blower the article but what you see below is what we call supporting evidence okay we're not claiming that which is discovered it which is discover this and by the way here's a supporting evidence for that discovery in databases called provenance because we can do this then we can do evolution of the story and I'll show that and then we have the pivots so in Wikipedia you have this see also those are the pivots remember when I was showing you the social car we have related queries and related hashtags when I use that to pivot on the story and we can archive basically we can archive the evolution of the story and at the bottom when you see references again similar to Wikipedia those are all the sources that we use so we're also going to measure how diverse our sources are compared to a regular Wikipedia page alright so how we do this we're going to harvest entire human data using two extensions one is called pseudo socials to the relevance feedback so pseudo relevance feedback is a well-known IR technique we're going to expand that the same with query expansion we call it social query expansion and then at the end we have a wiki fication algorithm so to summarize we read the tool firehose then for a given topic say us selection or brexit or what have you we have we're going to retrieve all those links by using s PRF and sqe we're going to rewrite those links and on top of that we're gonna build the wiki like audio so how does it look like the first one is we cannot rely on sis retweet because those counts are not good enough they have the they can artificially inflate popularity so we're going to build a new data structure for four counts and this is imagine this kind of a record we're going to compute counts at different level different granularities so there's account for engrams there's account for hashtags the second account for link those are counts at the element level then there are accounts at connections so if a hashtag is mentioned with a link we're going to count that if an Engram is mentioned with the link we're also going to count them and then we'll do counts of frequencies the number of times appears in the data as tweet retweet total and then both the number of time the users who posted retweets are tweets and total anyway the data structure is on the right and it's basically a record that will that will count different things and we're going to use this for different manipulations and different expansions the second part is we're going to use a hashtag index so we're going to build a table of cheese index but our relevant hashtags and links and the connections between a hashtag and all the associated information we're going to put all the social signatures they're all the contextual vectors stuff the ready saw and we're going to sprinkle the feet Twitter's provenance so it'd look like this so here's an example of the hashtag for hashed recount 2016 on the left you have different entries so December the 2nd December the 2nd December the 2nd December 3rd December 3rd as you can see the vector you know the first one is just ein Donald Trump Michigan recount and as you go down starts to change a little bit because it's expected recount the meaning of recount for December the second was slightly different from December the 3rd and it will be different say in 10 days the same with the social signatures the same with a related hashtags so the first one is basically Wisconsin at the end we're talking about Pennsylvania and then we have all the URLs ok so you can already imagine that you can use this hashtag to start seeing how the story evolves so if you want to know how the event recount happening you can just you can you can think of like hey let's have an index of all these things and the query over time and build a story the second is ok that's that's very cool has all these things but some of the things are duplicate some of them might not be good etc etc some of them might be only useful for a particular context in time so we're not use this technique called pseudo relevance feedback so the randomize feedback basically you retrieve of set of documents and basically the documents that you read 3 we're going to take a subset of those say titles or abstract and reissue a query so I'm going to do the same idea of social pseudo relevance feedback but instead of using the doc the content of the document we're going to use the hashtags all right so I'm going to rewrite the query rewrite the query with the contextual vectors so here when you say recount on that day means wisconsin' recount if you say recount on December 3rd is basically Pennsylvania Rica that's the cool thing right so if you are fans of war embeddings it's kind of a cheap way of doing embeddings ok but it's very scalable and it's only valid for this particular day which is what we want to do so the initial retrieval we're going to retrieve all the links associated with the hashtag or entity so those signatures work for entities hashtags etc we're going to rewrite the query using the contextual vectors and then we reuse the vote all the votes that we accumulated for weighting those links and then do a final rear rank so we have all these different links associated with recount and we have a data structure that contains all these different boats granularity we're going to use all this to pick the best links for that particular day and the other the other part that we're going to do is we're going to find similar hashtags so here's an example of sorry I told you that US elections to list exceeds the best data set because we've got so much fun looking at this on July 31st for the Corey Clinton you can return Hillary Clinton and that's a contextual vector but the really hashtag for that day is basically Benghazi on September 27 same query Clinton is Clinton email scandal and the hashtag is call me resignation okay that's Democrats let's look at Republicans okay August the fourth is from Train Mogga make America great again September 16 Trump nation Trump army the idea of the of these similar hashtags is you can query or you can expand by all these different things October the second Bernie or bust you know the hash tag the most associated hashtags hash Millennials November 11th shouldn't be burning steel Sanders the idea is these two things have the same meaning although the hashtags are completely different I mean unless you want to measure if they share any characters but they're there syntactically different they have the same meaning and on October 30th rogue one that was basically Star Wars and on 2016 is because one of the actresses and main characters of the Star Wars series passed away then the hashtag is rest in peace princess alright so that's similar hashtags once we have the hashtag index and all these different ways of doing expansion the last bit of the K is basically the evolution of the story so we're going to start with the hashtag index when we loop through all the hashtags over time we're going to expanse we're going to do these expansions on a daily basis it's kind of a temporal and dynamic expansions on a day we expand what we have the next day we expand or whatever is on that day and then will rewrite the links according to social pseudo relevance feedback if the link is not present we added to the timeline and if the link is already being discovered then we don't we don't add anything else we pick another link and then because we have all these rich data we can automatically generate all the queries which means that not only we build you a story the story already has embedded all the queries that should fire the story that's pretty cool right so if you're looking for you know Bernie resignation we'll have that story if you're looking for common recognition you have that story and looks like this that's the u.s. elections I know it's kind of a bit blurry but on the left you have a table of contents that is done the traditional hit list clustering on the middle you have the story that's the timeline I go that I just described as you can see is not they're not tweets all the data is coming from tweets but those are not tweets okay so what we show is the image of the article the blur or the snippet and then supporting evidence in this case there's an entry on Trump and then there's Fox News the that account share that particular link so that is the supporting evidence the pivots are the related stories and those are going to be the related hashtags and the topics and then for all the sources all those are the references those are the domains this is cool because we're gonna use you know Fox News political CNN New York Times Washington Post whatever has been heavily shared on social media will be there and all the queries are derived from the story as well as the annotations so in a nutshell this is a new data set that not only we can build this automatically but it's already and has already embedded all the queries that should trigger and has - all the annotations we say kind of standalone data asset so how do we evaluate this the first one is we do an offline evaluation on defect rate so how bad is the hash tag how bad is this entity how bad is this entry on the time line sorry Oh offline means I will show a blurb of the story in front of judges like in a Mechanical Turk scenario and they will ask them to find it there's an e also the question was yes yeah the question was how what do you mean by offline evaluation I was trying to explain how we build a hit that given a particular piece of the rectification story they can check if the entities incorrect if there's an incorrect entry query hash tag etc so they give us a an okay defect rate 5.6 percent you know so automatically we have some issues with some defect and then we're going to measure the diversity of the domains so we're going to take the Simson index which is a well-known diversity measure from biology and then we're going to compare our D index with the Wikipedia the index so if I go to Wikipedia at the end they have all those references so we go up we go at the bottom of our page as well for the US elections we get all our references we take all the Wikipedia references and then we compute the seams of the index on both data set and and we win we have we're more diverse than Wikipedia then we did comparison a temporal comparison against Wikipedia now we have a few Wikipedia pages will contain timelines so they don't have a volution of the story just have at the end what happened and they have there are a lot of wikipedia pages that contain the needs update label in any case so we compare the references that's what we wanted to know of our timelines versus Wikipedia we measure precision recall here's an example of the hash tag Harvey Weinstein and the other one is Manchester Arena and the coverage birthday is very similar to Wikipedia this shouldn't be a surprise because it's also human sensing and our stories usually tend to break much earlier than Wikipedia which also makes sense because Twitter is usually breaking news and they will take some time for someone to write an article on Wikipedia so what we're trying to say is that we can build this page currently that Wikipedia we can even say you can even think of like we build the page earlier and then say Wikipedia Wikipedia and we'll take over and do the addition there will be another thing and if we take toc table of contents comparisons for Sonoma fires ours versus the Wikipedia article is basically the same right so on the right that's your official Wikipedia article the time is October the 8th and nine ten eleven twelve and then from 13 to 31 because probably they will keep ears ago tire on the left is exactly the same the eight nine eleven twelve thirteen fourteen etc and here's an entry for October the 10th here an article of the fire track NBC Bay Area talking about the the fires in Napa and then you have a lot of different people so Napa fire wine Sonoma we also have very generic terms that sometimes are not good speed notes like breaking for example we know that's a problem and then the references are coming from all over the place from Mercury News Google San Francisco Chronicle CN CNBC SFGate etcetera etcetera so it's very very diverse alright now that we have all this the rest which is basically compiling a knowledge graph the the four components of the knowledge graph are what we call by the way lightweight knowledge graph so this is not a knowledge graph of all the knowledge of the planet but just basically what was said on a particular domain we have the links we have the topics we have the entities meaning people organizations places and we have time because we dump snapshot of all these things so we're going to look at the graph from those from that view so given a link for example the New York Times article then we can see who are the topics associated with this link who are the entities associated with this link and what was what was the annotations and relationships say on October 10th versus October 11th or November the 10th this year or the previous year so here these core schema if you're familiar with Excel pivot tables imagine doing this imagine you read Twitter and then it just basically run peanuts so you open another another excel sheet and say just grouping by users grouping by posts or grouping by links or grouping by by topics but the cool thing is once you build all these groups and all these aggregations you can traverse the graph which is you know conceptually very simple given a leg say we can we can see all the things that he share all the posts and all the topics given the topic Mogga we can see all the links about Magga all the people or entities associated with Magga all the links of the day with Magha and so forth so there's a plenty details the papers are also more details on the main tables I can't go into the other things unfortunately but you can imagine that once you have all these the rest is kind of straightforward I think I'm almost done opening for questions I just wanted to show you three examples of aggregations out of social data the first one was how to search for things to do using location-based social networks the second was how to build a search engine for the social web if we only use links and annotations and if you do that you can build a Wayback Machine which is something that could be super : and useful for other people using the components from the previous project we build a knowledge graph under the covers the powers the generation of a wiki light document for the construction of the story internally we call this storage TBS so its database of stories the shared components for the graph are basically the contextual vectors the social signatures the sauce the social sewer world and feedback the social query expansions and the voting their structure those are part of what we call skg social knowledge graph and we believe the techniques are generic enough so you can use this and say build the same at Facebook LinkedIn reddit what-have-you thank you so much for your time I know it's late and I know it's rainy so I'm going to hang around happy to answer any questions and ask me anything you want that's my Microsoft email address and that's my Twitter handle thank you [Applause] short me let me stop this I'm going to represent because I knew that was coming and I wasn't not going to present this but I have an example of fake news so in a second please long story short yes we do we do detect social sorry a fake links yeah I know let me just put another slide deck and just give me a sec it was intentional right so you saw the the question is do you guys care about fake news and I said yes and have you done anything with that we've done something so remember that I show you this pipeline where we discover good content from good accounts and good links so we tested our pipelines so this is our pipeline and I show you how to build the wayback machine and all those things so we're going to use our pipeline I'm gonna see if our pipeline can detect not a fake news a fake link again it's a fake link so for that we used this data set it's a fake news case study it's a BuzzFeed article has three time buckets so there are two real 23 links and 20 fake links a fake link will be for example you know I don't know this is fake or now but say that put in decided to intervene in the US elections people believe that and let's assume that is a fake it's a fake news yes but just what is on the other end this way correct so we cannot we don't know this the story is fake or not or our goal is to see if we can detect that fake link in in the pipeline early on okay so for doing that there's this data set available so you can download and this is fake links on Facebook our pipeline is on Twitter but anyway we'll give it a shot so gonna process the data through our pipeline and we want to know how we how we handle fake news so a bit of difference from the original study so that is Facebook ours is Twitter so that's kind of the main thing that the total number of share is higher on real versus fake so the original fake study it says that as you get closer to the election day there's more fake news the number of fake news is higher than the number of mainstream news we don't see that we see that as we get closer to the election day both there's more real news and there's also more fake things okay but this is kind of the the the cool thing how we're going to process the link so we're going to start in four steps so see step zero is we just kickoff by the way to clarify there are three buckets February April May July and then August election day so February April according to the data set is no more fake news as you get closer to the election day goes higher early on high good news and as you get closer to the election day there's more fake than real news so we have four steps in our processing the contents election remember how to select your content how to select good users how to select good links and then the final cat we call it that way those are step from one to four so here the blue is real and the orange is fake so if we look at February and April we cut basically the fake links in step number three so in step number three we were able to get rid of most of the fake links and then by step four they're not fake links although the number of good links you know lowers a little bit if you go to major lie pretty much the same know by step three we get rid of most of the fake links and step four all of them are gone although we lose some of the good links and as we get closer to election date although we do remove fake links we still have a few fake links so most of the fake links are not viral enough to be selected so we know they don't populate through our structure and out there we win them out early on now if we get closer to the election it gets much harder to filter out fake links this could be because you may be a trusted user session and sharing a fake link and still the step number four reduces the number of fake links pretty good so we believe that with the pipeline that I presented if you're in the business of the techniques a fake content you can do this early on even at the link detection without even spending resources on the technique the story is fake you can already cut this link early on at least you can label as a potential fake link you can fine-tune parameters run more optimizations for example for the content selection we can say we need 200 shares per link so we're going to put a video of a threshold for users so we want one verified user and five users from trusted users added up to ring - so verify Bill Gates myself ring one Alex a ring to unlink for example four for the link selection we can play with some some of our variety scores and even if we do that we still get a few links at the end but we remove the fake links way we're early on so much earlier than here so in step two most of the failings are gone so if we want to do this at scale again saves you a lot of computation time we don't do this on Facebook because we don't have access to Facebook I don't work on the spam detection team so it cannot come but we do are in the business of detecting fake laying spam links etc any other questions yes sir we I don't know today but we used to pay for the firehose so uh man and you should cross the line to evade those lights with the Google does they just crawl static babies I mean Twitter pages so this project this project is based on firehouse access so we at least at a point in time we do have access to we did have access to the to a fire hose everything is assuming that we you have access to the tooter fire hose so the question is do you cover only news which is in for the two projects are mostly about human sensing human sense I mean you either tweet or write a post but you also have to do a lot of sharing regardless if this is news articles or a video that is going you know with dogs and cats which is only care about large-scale human sensor I hope that answers the question it's just not about specific domains about large-scale human sensing yes follow-up question runs every day well the thing let me see if I have this I'll get back to it in a second the the time the social knowledge graph which is basically four components links topics entities on a day but every day used you timestamp this graph so it's not a single graph so every day there is a snapshot over the graph but because it's snapshot at then you can go back in time on Bill dugg regressions that's how we build story evolution that's how can build this Wikipedia page that's how can we build the Wayback Machine so it's not a single graph so there's no like if a graph is every day there's one snapshot maybe today there's less activity than yesterday because maybe people were talking more sharing more yesterday than today or maybe tomorrow but that's the thing every day you dump a copy and then after that you build our creation so you can build say that we give you the monthly view the year view the decade we radiation after to be we don't so the question is do you guys do deltas or accumulate there this is kind of a building block so every day it is a snapshot if you want to build something because it's an internal graph you have to kind of parse whatever part of the graph you want and bet and then build whatever aggregation you want to build on top of it so if you need a delta you have to build that Delta if you need to say another aggregation on top of it you have to do it I just described the the bare raw be autograph [Music] savior's does the hashtags give me the meditatively watch it drift because body building static pages off of the change of some fear many hashtags attacks the descriptive tags that come off the tags you Auto kind of build a static you got stuck in this you also have dynamic want to show the drifter to change you could you could I'm not saying that you cannot I'm just saying the the the way that we've done it is we kind of download a copy okay of what we believe is the knowledge graph for this for today tomorrow is different yesterday was also different correct everything yeah yes what would it be and okay correct so the example here is why we're dropped if I can understand the question why we're dropping some of the good data although were using trusted users [Music] Oh content goes first because we want to see a it's just the order of that's a good observation it is the order of how we build a pipeline because the content selection here is all the tweets we cannot you know we don't want to process and run all these things in for example two hundred and thirty million tweets we want a subset of that okay so think of like you know tweets links user hashtags and engrams and every in every face from content selection links user link and final cut we always kind of get less and less with the assumption that is good yes will will law we're going to lose a few things may be good things but it's just it's just the way we're building so the question if we do more complex stuff to infer of different things about the world know this at least in this project no it's very very simple in the sense like think of this as like a massive filtering at scale once we have this in place which you know took us some time to build the rest you can run all those things you can do you know fact-checking and a bunch of other stuff I don't have anything to show the moment on that oh that's for the things to do yes sorry I thought you were talking about the we could be that can you repeat the question again for certain facts okay now for the things to do is mostly location expansions okay so we're just using some formulas from math oh yeah that kind of stuff so the fact that a P oh I you know is in a radius of certain miles or kilometers then we'll just derive another city or another place state county country that kind of stuff I know they don't mind of yeah it's mostly geolocation no no no the question is its most about geo containment okay so we don't derive any the question if we derive anything specific no we don't it's mostly geo containment Alexi of all the ugly little eyes so like how do you make sure that which is similar in raps there people it's all about volume so let me see if I have the the devoting the restrictor here is small seed about volume so when I say normalize it's mostly normalization that they are clean they don't have any problems and then we just count all the sorry I don't have it it was on the previous one I know it's here which is basically normalize all the engrams like just a tournament token and then we graph for that particular thing everything that we can know about the term so if the term is Trump it's a trend but also is an entity Clinton it's it's an Engram but it's also an entity we also do other things like it's not here cash tax so dollar MSFT will be Microsoft in terms of financial information pursues hash tag emissivity she's basically kind of taken the hundred forty characters and kind of fine-grain accounts on different elements of the tweets basically at different levels because we do currents then we know that hashtag mentioned with the link link mentioned with the entity link mention with the Engram and so forth and then counting good for example yeah question are the and yes well we use it this is a research ship yeah this is a research prototype part of the graph is there on the spotlight shego and a big new sensor will be part part of the stuff is power in the rundown the rest has been mostly a prototype on how to build knowledge graphs for internal consumption yes it's basically it's another data asset that can power other microsoft properties so I know when we show these people say oh you know build Wikipedia no no we're not building Wikipedia we're not externalizing any of those pages could be cool but it was mostly for us so we can understand how good the content was and if which one of the problems in in part of the graph or placed in by hand stuff help me help to to so it's a really good approach music happens especially in an area where we don't know what the value it is it's an actual open-source project or a project that's going open to our students the over whether that's called Bridget Nile they were to look at back into let's movie put their own opposition or support of the actresses from distributing things [Music] it's somebody makes money okay all right I'll take a look yes so the evaluation is very very hard okay and so the question is how should evaluate this this type of assets that's kind of a generic question how you evaluate the main specific graphs I don't have a good answer to that the only thing I can tell you is that the only way that we do the evaluations is on specific cuts of the graph so you want to evaluate the users you may evaluate the links you may want to evaluate the topics so just showing the graph in front of people is like what the heck is this so what you do is you query the graph you get some data and then you build an evaluation on that particular angle of the graph and then if you have many many nodes we'll just pick your notes and then you iterate so you know that for example we're very good at users but we got some promising for example in this X in this project we have some problem with certain hashtags like I show you break in yeah it's a very popular hashtag but everybody's breaking harsh news not really good so we know that we have some problems there in users we don't have any problems in links were pretty good but in certain hashtags we do have problems and the only way that we found that out is because which is evaluated on I didn't by item on the graph now obviously this won't scale if you have like very complex knowledge graph but this one are what we call lightweight knowledge graphs for five main entry points so then it's more manageable but it's a hard problem I yeah I don't have answer not random in the baby or the things to do an exterior is what is it is simply like this and Medical Sciences it means to be kids but then things go there is going to be pretty sure for them do we bring that up on the wall sorry for changing but yes hold on because I don't have think I have things to do only on this one and so give me a sec so you want to know why we change things to do okay I'm presenting again I'm getting good at this okay so the question is secures things to do in San Francisco and the first one is Golden Gate and I think you're saying why we're not showing Golden Gate right well the reason is because all these the things that you see on the bottom those are POS that have annotations that are for kids okay so that's what we built maybe I was going too fast when we build this classifier we have different aspects so you got all these different constraints on the taxonomy and then you have the POS so if the Academy of Science has a lot of tips about kids okay and then you're looking for things to do in San Francisco with kids if you didn't say kids you know Golden Gate goes up but because you say kids you've got Academy assign was number 20 now is on top right and the way to do there is to leverage the POS with the annotations annotations are the tips that's what we do we can do romantic things to do what I can obviously we don't have all the aspects because some of them don't have a lot of data but that's the point so you see things to do in San Francisco is one recommendation imagine with kids is a filter on top of that recommendation okay and all those things do have a tip that says good for kids good for small kids lovely they with a kid all that kind of stuff yes that's that's the and this is for the other gentleman this is geo containment is basically moving up moving up and it works for Europe as well so there's a CIA paper and then there's a demo I see that when you assign where we show all this all right oh you got a question yes [Music] and also how far you think business well the first one which is things to do is in production so this I show you this already in production for the fake links that data asset is useful Microsoft's mostly to identify hashtags ready hashtags links the Wikipedia example it's just an internal demo that we put together but the real the value is on the graph that powers the applications which is so in an example that if you build this knowledge graph then these are two examples so the the wayback machine that's an example searching the social web with the social cards that's another example all these are examples of this this is the intern this is a search in Twitter from a different perspective if you have a graph okay so maybe you would like to see examples from Facebook or Twitter or Pinterest this way we're not claiming they have to do it that way but here so what we done is basically we read a social network data and then we spit out a graph if you spit out a graph then you can build these kind of applications on top that's what I'm saying moments are literally driven so moments are my understanding is literally driven so people someone make sure that the moments are okay this is fully automatic everything everything I will show you here is completely unsupervised feel free to take the pictures yeah that's an websites paper couple years ago yes here one question users first so the idea of trusted users is the following so you want to imagine you want to read say Twitter you want to read content from quote quote good people okay so how do you know there are good people meaning good accounts one way is to go on only read content from verified users with the gate where a coab AMA Donald Trump etc that owned itself is already bias all right and will have a lot of companies you know CNN you know Microsoft they're going to be there if you want to get like average Joe like regular people you have to scale and one way to go is to you this technique that we call trusted users and we can go to ring number ten and that's what we stopped there are many other ways of doing this it's just one technique one design but the goal is no we went out to attempt by because we we couldn't find more better content after that no no deep learning here this is nothing against deep learning this is no there's there's no deep learning here this is mostly heavy on information retrieval data mining and also a mill machine learning is mostly for the detection of good content so no it is a it's a it's a machine learning model so the ML Molly will tell you the content is good this is support vector machine if I remember correctly we use a lot of different features like who is a user number of hashtags ratio of hashtags versus regular terms presence or absence of a link etcetera etcetera number of ads to some people I can point you to do has all the details of the classifier but basically the ML model is the selection of the content this you know you can if you want you can build a machine learning models but this is more of an activation Network okay you start with the seed again could be labels like super good users and then voila you scale the viral link detection this is mostly about kind of looking how the link goes through different viral structures within the within the graph and then we're trying to pick links are heavily richer it's most about activity and the other one was the vectors you mentioned embedding it's kind of a cheap way of doing embeddings okay the reason why we didn't do embeddings in general is because these are super fast to complete search this basically counters and the same technique works for the hashtags and also works for the links we don't have to train without to retrain there's nothing it just like it's basically a grep and then signatures also uses a model models on the on the graph on the hashtags the topics or the links the AML model for the the question is how long does it take to train we don't train the ML model for the content is done so we basically every day we build the graph so it takes a few hours to build yes we did labor our own data that was extremely painful we can talk I you know I gave the talk the data by the way a few years ago how difficult is to label to it in general how difficult is to label social data so if you're interested we we came up with a methodology and how to label complex datasets okay so the example is I show you hundred forty characters and you have to label this and it's very very complex however if you can label some of them then you can choose you know build your straightforward machine learning model but the key is how to label social data which is hard living facebook is hard living P wise the location-based social network is hard labeling your LinkedIn profile is gonna be hard cuz you're the only one who can label that etc etc so labeling social data in general is very hard and if someone tells you that they have a super cool high interior agreement label little data data set from social don't believe them because it's very hard very very hard I can point you to a few examples yes sir yes expertise we did something on expertise I don't have anything to show here but we we did and nothing we have something on how to expert this detection from social networks so the way of how expertise detection is done is the number of times you mentioned a topic or a hashtag that's your expertise if you mention a lot deep learning they know you're an expert on deep learning so we extended that model and then we build another thing which is not only what you're tweeting or writing say LinkedIn or Facebook but what people are tagging with you so if people are tagging you with certain topics that is also included as part of an expertise model i fortunately I don't have anything to show today but if we're gonna talk offline yes we've done something to show you the poster the area so we used to have access to cloud data like six years ago I believe Microsoft was an investor or something the problem with clout I remember we have certain things are kind of off I remember tea bowl you know the nf former NFL player was tagged with things that were not appropriate for football so it's a little bit hard to gather expertise on that so we look at it and we decided to do something a little bit different but we did we didn't look at cloud data for quite some time yeah but but the Geo the the percentage of geo tags which is kind of small so yeah any more questions if not I'll be around until 8:50 cuz my parking ticket is it's 51 so have one minute to go on grab my car but I'm here for 48 minutes yes she's very over today and all we need volume alright so say that you have what the hashtag xyc was very popular today and then disappears okay this presentation is mostly about volume however our graph you can also build what we call domain-specific graphs so then you can lower the threshold and then you can get stuck there's no variable has a lot of volume but I will gather these things like X Y C is consistent over time which is basically what you want this particular presentation is mostly on high volume but the techniques work for that you can choose its particular configuration file yes so it doesn't show here but I yeah so sorry let me just go one by one see if you're asking show up here so this is one of the papers web science they're all on the web if not shoot me an email send you the copies web science then we have the activation at work is awsm the story evolution we have two papers on story evolution the JC DL and ie WSM last year I have an old publication on that based on the Apache CVS noise many many years ago we don't have anything new that publish I mean we don't work but can say anything at the moment all right thanks for coming [Applause]