Devreal

Omar Alonzo, Crowdsourcing, UC Berkeley Soc167 By the Bay

Omar Alonzo, Crowdsourcing, UC Berkeley Soc167 By the Bay

Recording: Omar Alonzo, Crowdsourcing, UC Berkeley Soc167 By the Bay

foreign I work for an old company called Microsoft I'm an applied researcher scientist whatever you want to call it I work on a product team by the way I don't work on MSR so I work on Microsoft research I work on the product team and today I'm going to give you a high level overview of this thing called human computation crowdsourcing Etc I have quite a bit of slides just to try to keep you engaged and awake in case I get into travel here's a disclaimer it's just my own opinions so crowdsourcing is basically hot you know there's a lot of activity there's a lot of interest in the research Community if you go to any conference you'll see papers on C IR knackle wisdom Kai rexis and vldb which is the database conference so when the topic gets to a database conference then it's super mainstream because usually the the last guy is to actually get into this there's even a dedicated conference called agecom human computation that was actually held last week in Pittsburgh for only uh human computation and a year ago and the past year before there was an industrial Gathering called crowdconf sponsored by cloudflower cloudflower they didn't do it this year a lot of companies using crowdsourcing Big Data based on crowdsourcing startups and VCS putting money into this so crowdsourcing is a book by Jeff Howard like eight years ago where he introduced the notion of crowdsourcing you can read the definition on top but basically it's the application of Open Source principles to feel outside software and probably the most famous example is Wikipedia Wikipedia is cloud source can you identify this thing captcha what is this this is a program written by a computer that will generate a test that the machine cannot pass okay so it's like a tuning test that machine cannot pass so you have two terms here one is the term that the system is trying to ask you to guess right and to make sure that you're not a robot there's going to be another control term where they only know the answer and company invented this because there was a lot of activity open accounts in email automatically by bots so the only way to stop opening those emails accounts automatically was to have a human intervention saying how do I know you're a human and not a robot okay so if you know this you are going to identify a few things as the UC Berkeley page in Facebook okay that's a place how to get coffee in Brooklyn Yelp we have a question on quora about Berkeley and on top of that you have a book from Amazon about node Excel you're going to notice a few things they have stars they have likes etc etc all that is doing by is done by what humans not by single machine okay so human computation is not a new idea it's actually quite all you are a computer I'm also a computer this is a picture from the World War II area where a lot of people used to do manual work that will get into building logarithmic tables and a bunch of other stuff it's a great book by David Alan Grier on human computation and you can see the connection of people doing things by hand and then the birth of the computer which is the automatic computer some definitions to make clear we're on the same page so human computation is a computation that is performed by a human okay so if you can do something the machine can also do something and vice versa human computation system is a system that organized humans effort to carry out computation and crowdsourcing which is the topic of tonight is a tool that a human computation system can use to distribute tasks all right if you want to read more there's this book by Edith law and Louis Vaughn Anne great book very small book that you can get a pretty good idea what human computation is Luis professor at CMU he's the guy who invented recapture he's also the guy who invented the ESP game which is a game that was used to label images and the way we label the images is you're going to play together this game and then if we agree on the labels for the image then we pass to the next round and they found that just because the game was so addictive you can automatically tag lots of image very very quickly there's another project this latest project called Duolingo so if you want to learn Spanish Portuguese Chinese you can go and buy Rosetta Stone but you can go and use Duolingo so Duolingo is going to be it's going to be step by step how to learn these things playing a game why we care about this because data is King it's the most precious thing today on Earth if you work on software it's a new oil right so this is a nice chart a few years ago you got a couple of references here where they show that you can have a lot of different algorithms for doing all sorts of things but as you start to grow in the size of the data set the Algos are starting to more or less you know perform at the same level so what we need is basically to collect different ways better ways to collect data and so far the traditional way of collecting data for you people in the audience mainly you know social studies is you set up you know some experiment you're going to recruit participants you're going to pay five dollars you're going to put a fly outside come to the lab do this for me and um you're going to you know do some metrics produce report and so forth that will take you an incredible amount of time and effort because it doesn't scale it's very slow it's expensive and you know you have a lot of bias because if you're going to recruit your friends your friends are going to say this is awesome because you're your friends they're not going to tell you the truth so a few years ago a team at Stanford started using Mechanical Turk which is one of the most important platforms today to run a few experiments about natural language processing so they wrote this nice paper called cheap and fast but it is good valuation non-experts annotations for natural language tasks the first author is that Twitter Brendan O'Connor is at UMass the last author Andrew and J is Mr deep learning now what Baidu before that Google and still at Stanford so what I did is they took five very difficult tasks from natural language processing and for 26 dollars they got 22 000 labels unbelievable and they found that the average worker aggregated over five workers per label were as good as the expert so that means that in this case they were able to get the same level of quality but instead of using expert instead of hiring those very expensive experts you can do it using Mechanical Turk so cheap and fast it was another cheap and fast by another Stanford that then was a professor on the East Coast tackling a more difficult problem machine translation we're not talking about translating from English to Spanish or Italian this is from Arabic to English very very difficult so this is an example I know if you can read this but I can leave the slider this is kind of different dialects in Africa and other stuff around Arabic if you think NLP is difficult Mt machine translation is even more more difficult in particular because it's very difficult to find those experts to do to create this goal set for you they paid 10 cents per sentence on the order to get awesome stuff also a few years ago this is not a startup in San Francisco by the way who sells food but um so Michael Bernstein who now is a professor at Stanford they build this very very cool system called Soylent so you guys are going to write a paper for the midterm for for the probably end of the year project another class project and since you guys are all Mac users are going to use something called war or something very similar and if you don't type very well you're going to have this kind of scribble reddish thing saying there's a mistake there's a misspelling so you can choose rely on the spell checker or you can ask the crowd to help you on that so basically what Michael and Company did they they basically build a plugin on top of world and it's going to help you write your paper term so you say can you can you just shorten the paragraph please because you know the maximum page link is stand I'm like 10.2 all right so they're going to take that create a Mechanical Turk Outsource it to Mechanical Turk collect the data and say whatever has been shortened and can you proof check can you spell check my document please so you're gonna every every term that it your speller is not detecting is going to out to him Turk the the crowd is going to correct you um you know it's going to be part of word back super cool super super cool so the three things these guys have in common um the NLP machine translation and Soylent is they leverage Mechanical Turk mechanical turkey is a site by Amazon um it's kind of a punt because in the 80s or sort of early before that when they coined the term artificial intelligence you know intelligence is natural so artificial intelligence is obviously artificial so if you negate that artificial artificial intelligence you get just you know people so basically it's a marketplace where anybody can go and build a task you collect your data you pay a few cents and you're you're done that's basically it it's not the only one you know Alexa already mentioned craft flower startup on the bay area and they use more channels so Mechanical Turk is one of the many channels but this kind of goes across the world so if you say hey I need to translate Chinese or I need someone who knows German etc etc this can help you for now you're probably looking at me and say what the heck this guy is talking about so what what is what can you do in mechanical Turks so we're going to look at an example of a hit hit stands for human intelligent task is the minimum amount of unit work unit that I can do so this is something that I just got a few hours ago it's a hitomechanical Turk that is asking you you know given a Trader Joe's receipt you know can you just you know tell me you know that there's an air salad spinach 199 okay so you read that you say salad spinach praise you know 199 you basically transcribe receive into a database so this probably go into your database because someone is looking for parsing this data awesome do something completely different a product description something is in somebody's into Harry Potter and I need to help categorize this item I have no clue about Harry Potter so I'm probably going to skip this but if you're into it you can say well I can I know what this is It's all about and here's the information you want so this is good few years ago I started playing with mechanical Turf for doing non-trivial stuff so this is a query where to go on vacation so I work for Bing by the way but probably you'll be using Google so if you go to Google and type where to go the autocomplete will tell you vacations the means that it's a very popular query now if you try Bing or you try Google the answers are actually not very good because they just show you a bunch of web pages about vacations and you have to read the title read a snippet open the page get an idea and they don't tell you where to go on vacation I just want a list of places to go on vacations so what I did is I put a task on Mechanical Turk I pay for 50 answers 50 answers I want a list of top you know Place top 50 places to go on vacations a dollar ninety cents I also pose a question on quora I got two answers by the way you can just check because you see my name you'll see everything there Yahoo answers to and then also post it on my Facebook page and my friends one answer and his answer was go to Abu Dhabi very very useful however the crowd give me lots of cool stuff so on the right on the left you see the list of countries so I have UK you know UA turkey Switzerland Spain Singapore Portugal blah blah blah us and India at the bottom so like a lot of answers for the US in terms of cities I have pretty much everything that I want so uh if you go if you go to the Bottom by popularity we have you know Las Vegas Hawaii um you know all the fancy places to go but there are a lot of like interesting places that I wasn't very aware that I can go on vacation so for a few minutes 1.80 not bad if you take it further if I can try the famous probability test you know if I flip a coin it's head or tails is that 50 50. so what I do is I pay Mechanical Turk for someone to answer me two questions you know pick a coin that you have in your pocket if you have a coin I don't care what is a coin could be a dollar a year or any other type of coin flip it and tell me if you got a hair or if you had a tail it's not 50 50 but hey you know why this is interesting this is interesting because if you want to collect data and you're going to be collecting data for all your projects and in particular in real life in industry you have to design experiments to collect this is very cheap and extremely fast you don't need to set up any infrastructure it's already done for you if you're building things you can introduce experimentation early on the life cycle and for new ideas very very useful however a number of caveats and clarifications so trust and reliability super important here wisdom of a crowd revisit you probably have heard this term like oh the crowd is always right the crowd is always right guess what not always the crowd is right you have to adjust the expectations of what you're going to get back because crowdsourcing is choose another data point for an analysis and that's all it's not going to replace anything it's just another data point and it's complementary to any other experiments you're going to learn as undergrads or grads why now because of the web okay so imagine the same way you have the web as you think they're just machines the web it could be machines and people the same way you think you have a cluster of computers you can have a crowd of people and each of them can do something for you um and why it's important because you can you know solve problems that computers are not good yet so we can use humus to collect data to train machines to do something that it was impossible to do it before and we have scale and Rich by the way a little Interruption I usually speak talk fast so if you think I'm like going way too fast or you're getting bored just let me know and I'll slow down who are these guys who are the workers well you know you can imagine it's not someone in Midwest you know a wife being bored say oh what can I do I'll just do some M turkey sometimes that's possible a lot of people have different um reasons for for doing this kind of work sometimes you just need the money sometimes if you're International you're learning English so this will be really helpful for you to learn English better sometimes in the case of people who are like really for example shopping a lot they shop a lot in Amazon they can get like you know extra sense that maybe in the few dollars you can you know buy an extra book an extra CD or another download um so people have studied a bit kind of the phases of mechanical Terror how are these guys etc etc however I want to uh spend a bit of time on the Dark Side of crowdsourcing this is an article that appear uh like a week and a half ago in uh where my Wire magazine talking about um well the title says it all right so I'm sure you guys are a hundred percent on Facebook and you spend like copious amount of time on Facebook right and you think Facebook is awesome you know I have a ton of friends working at Facebook but the curation of the images that shouldn't be on Facebook are done outside of the US here's a picture of a third-party vendor in Philippines where there's going to be a human computer a person looking at a picture and say looks old I don't know maybe all that data will get back and then you know learn a few other things so next time we don't you know they don't show you this so be aware that although human computation looks great and there's a lot of potential there's also a dark side and a lot of cases someone needs to do the work and unfortunately you have to be careful because not everybody will like to see certain type of pictures when you sign up to do a task all right so you you probably see like what's why are these labels why people are so you know into labels um I was reading the syllabus for the class and you know you have a machine learning part and you have a data science part all this is needed because there are two little things that are very important on a search engine assessments and labels and let's discuss them in detail relevance assessments are the key to make sure that your search engine is working all right so the the task given a query and given a document in this case the queries Milton Keynes the town in the UK and here's the the search results from Bing and the question is given the query Milton Keynes and here's the the list the serve the serp engine result page is this relevant or not okay someone is going to say it is a relevant document to the topic or might not be relevant okay those are assessments and are super useful um just bear with me for a second last night I took this screenshot from a retweet that Alexis Alexa did this is at your typical machine learning framework and is there any MLP is in the audience please forgive me but usually the workflow is you get some data you do some cleaning um there you have a lot of feature engineering you choose whatever ml method you go and voila here are the results you can visualize Etc do you see anything interesting in this picture this something that kind of just grabs your attention nothing you have a lot of things to the right but not a lot of things to the left okay so a lot of make the assumption that God gave you data it's not that's not the case couldn't be further from the truth so be very very careful with this in the area of a particular the area of big data and machine learning the model is you get some labels you're going to do some engineering to produce the features you're going to be the predictive model and then you're going to work on the optimization which is cool but the Lebanon experimentations are usually perceived as extremely boring but the advice will be do not rush the labels because there's people's people and machine involved here and stay with me for a second on this in this threat label quality is very very important you don't want to Outsource Outsource it and you have to own it end-to-end whatever you do even for like the simple things and if you work on large scale which I do this is super super important why because data Gathering is not a free lunch you really need to get this data you are you're going to set up infrastructure you're going to ask people to do something for you you need to Cather this direction not that temperature okay and the temperature maybe but these things you have to gather labels for the machines are not labels for the humans and there's an example in a few slides on this very very important there's a lot of emphasis today on the ml crowd on models and optimizations and Mining from the labels but not so much on algorithms for ensuring you know high quality labels and if you're going to build training set this is important okay enough mumbo jumbo and why again why the labels so work with me walk with me sorry on a little example of why this is so important for building a search engine so I'm going to give you a document I think I'm going to say could you please assess the relevancy of this document to this topic that's the first thing okay and you're going to produce you're going to judge the document to the query and you're going to produce an assessment or a judgment okay but that's me Alexa is going to do the same and you're going to do the same the three of us someone is going to collect all these things get a majority vote some some trick and it's going to say the by aggregation these three five ten people are saying that the label for this document is the following relevant or relevant or whatever this is a label okay so there's little thank you we're going to use the label to learn so you're going to learn something when we learn something you want to produce a model when we have a model then we can predict if we can predict then we can label the rest of the collection because you cannot build sorry you cannot label by hand entire web it's impossible what you do is you sample a subset you label that and then you turn the model to label the rest okay so you can never to do all this other labeling and the last part is if I if I have a model and have a label and I can predict I can evaluate how the quality of the search engine works you follow me more or less right again another way of looking at this is garbage in garbage out okay if I don't gather the right assessments and write labels I'm going to learn the wrong things I'm going to predict the incorrect things and I'm going to evaluate the wrong stuff that's not what you want all right Switching gears to information retrieval so information retrieval is basically a term for search engines in case you don't know about this and I have a friend of mine Mateo Birch who lives in the city and he did a few cartoons for me when I gave a tutorial a few years ago so this is a you know someone in the North Pole trying to assess the query snow in the North Pole yes you know it's a super relevant whatever so relevant judging is extremely difficult to assess because it's very subjective it's extremely expensive and if you go for if you if you talk to people like Microsoft or any other company who maintains the production search engine usually have to hire some professional editors or some sort of editorial work to help you that and the benefits of crowdsourcing the potential benefits are the scalability is going to be super fast and very cheaper to do and the diversity of the judgments it's just one like a sample of the real world telling you how this thing works so when my friend Mateo did my first cartoon I said this is cool can you do me another one on relevance assessment I don't know if you can actually read it um but basically is the teacher says you know Alexis says next I query for idiot and get back a photo of a reality television start and then you have to tell me if this is relevant or not relevant so I thought was an interesting joke so I implemented the joke this is a hit from Mechanical Turk so relevance uh questions say you query for idiot in the search engine and you get back a photo of a reality television start what do you think this is relevant or not relevant so the result for idiots are in 507 said it's very relevant and 207 so it's not relevant and people actually wrote me I can leave the slide egg with Alexa you can read this but people are very nasty on them you know television is one of you so this is the why it's very relevant to return television star when you issue the query idiot and two are saying well not exactly so still majority votes wins here all right so you have a new idea I have a new information retrieval technique but you know I'm at Berkeley I just don't have access to click data man I just have to obviously the query logs I don't have the money to hire editors how can I test my ideas um we're gonna do it so the cool thing is you don't need to come into the lab we're going to use diversity we're going to pay as little as possible and it's going to be super agile so you go back to your advisor and say hey I just read a bunch of papers and this is the way to go we're going to go crowdsourcing 100 um uh just a few you know hundred thousands later per month and you say easy so it's not going to be easy we're going to go through the details of getting this thing right the first one is asking the right questions so instructions are key in this case the workers are not information retrieval experts so do not assume the same understanding in terms of Technology so this is you have to write it in plain English it is always important to show examples and if you can always hire a technical writer or someone who has an English major so it can help you build up that I'm prepared to iterate over and over and over so I'm going to show you an example of how not to do things so this is a task it's basically saying search this is all entered by the way search for a topic and collect details about advertisers so they say go to this website okay go to the website click on it follow me in the menu in the menu on the right side you will find a menu entry search click on that menu entry which will take you to whatever page or go here okay seven items to remember um oh search for Mustang copy the URL enter the URL of the top Place Advertiser okay so 10 items for like Cent ain't happening a lot of work I'm not doing this so time to brush your user interface design class or book you need to grab attention generic tips okay the experiment should be self-contained you got to keep it short and simple clear with the task you have to engage with the worker and always always ask for feedback and if you are going to do it in multilingual you have to localize in English first get a ride before you move into German or Spanish or whatever second is how much I pay it's very important it's a very delicate balance so if you paid too little like one cent no interest if I pay it back guess what I'm gonna get all the spammers okay in particular all the robots so you can start with a number and see if you got any attention and if people are doing your task ball maybe it's right so but if it's getting slow maybe you're going to pay a little bit more or you can pay by effort so you can get a metric that says I'll just pay you know two cents per click whatever the click is if I had to answer two questions there's like two clicks then I'll pay for Cent and stuff like that you can also pay a bonus so if you do well there I want you to do more work for me so I'm going to give you a bonus which leads me to The Next Step which is managing crowds managing computers in the I.T Department doable managing crowds slightly more complicated why because we're gonna um put some quality control mechanisms mechanisms into this and the quality control mechanisms are the following we want to assess quality as the overall in the experiment so it's not just are the workers not doing their job but also maybe me myself as a requester I'm also very bad at doing my job so qualities overall overall you may think the worker is doing a bad job but maybe it's very sloppy lousy requesters um and when do we assess the quality short answer is before during and after so before you start the task you can put a qualification test the same way you're taking this class right you gotta you have hope hopefully you had a pre-qualification like you have to have this class before you're taking this class and this is basically to screen to make sure people know about the basic things then you can assess during very similar to captchas or very similar to uh you know goal hits like how to make sure people are paying attention and after once I'm done then I'm gonna without all the bad performers this will be the equivalent of a test and based on that I'm going to filter calibrate and just produce a report again if I'm going too fast you let me know please um how do we measure worker quality and why this is very important remember the ultimate goal is to get labels high quality labels right so we're going to compare the workers label versus a known untrusted answer so that's known as a gold hit or a goal set so someone create an oracle of things that are actually the truth and then we compare against the truth so if you pass then you're good to go if not you are not doing your job however a lot of cases it's very difficult to get a ground truth so you have to do something like compared to what the rest of your peer group is doing so if they're telling you please label the image and you're saying it's red and everywhere else is saying blue you're basically off okay there's something wrong with with your understanding of the task then you can build models to predict you know the worker performance and things like that you can also verify the workers label you can do that or you can try what is called a tier approach so the same way you know programming languages you know they have these patterns it's also a pattern called find fix verify that's the pattern that they would use in Soylent so the idea is I'm asking you to do something you give me the results then I'm going to use someone else to find there's a problem someone else is going to fix it and someone is going to verify that the correction is correct okay so everything in in multiple little hits but at the end you get basically what you want and then the other piece that is super important is measuring agreement so why agreement and you're going to see this over and over in a lot of the literature why agreement is important because we want an example of a search engine we want a label that says this is relevant or not relevant for the document if I pull five or six people I want to get agreement on this so there's like lots of way of measuring the agreement I will see what happens when you cannot measure agreement okay so the basic thing you have to do is measure the agreement between the Raiders so what's the agreement between ourselves if you got a goal set is what the agreement between the workers and the goal set and you can use statistics like coins Kappa which is well known from from the literature for two radius flights Kappa an extension for any number of operators crippenders Alpha for all sorts of things but you're getting into great areas like I have one document and I have five workers five judges to say relevant three say non-relevant is that good enough the maturity seems like a little bit weak so maybe we're going to get another label or do something else or maybe we're going to leave it as this for the entry-level workers but we're going to use a tier system so we're going to ask a super expert to break the ties yeah then content quality is also important because you want people to work on things that they like so say that um you're not into sports you don't want to just be judging content about sports so it's always important to refresh to randomize it's also important to have modern content track is an example it's a well-known collection in information retrieval they have a lot of the airport topic is pre-911. so it's kind of very dated you know you're not going to get anything very interesting there document length people don't want to read like pages of pages of stuff so just try to be succeeding to the point and then avoid working fatigue so if you do a lot of if you're not Turk you're going to do a lot of work and you just want to make sure that whatever you're doing is interesting and engaged you don't want to be like oh my god when can I get over this task like the speaker one is going to be done was the task difficult there's also another question and sometimes the task could be difficult so here's an example of assessing content on a wide range of topics and they thought that uh you know for airport Securities and everybody you know has flown a plane at least pretty easy to assess the relevance assessment but if I go into a Schengen agreement she's only the EU you just don't understand this you're not into Greek philosophy you just like don't care so these things are like getting more and more difficult to assess so not every every content every task is equally easy you know there's a different grades of difficulty all right we're going to take a little pause here and probably right now you're saying okay that's like super super cool but you know it looks like a ton of work it is sorry I have bad news it is but the original goal is data is King okay and the quality and experimental designs are preconditions to make sure you get the right stuff so do not cut Corners so hang on with me for the next uh 10 minutes will go and see how can we actually without all the bad stuff and actually get something cool um so ground sourcing works I mean it's you know faster around easy to experiment in a few dollars to test but as I say you have to design the experiment carefully um there's a lot of issues on the platforms a lot of issues on quality okay say that you know okay more or less I know how to deal that how to deal with that but if you're going to go in production systems when I say production imagine you know Facebook Google Yahoo either in Microsoft like lots of things going on you have large-scale data sets in the millions of billions um and you you execute these things every day it's not like a one time it's like hey I need 10 million labels I'm done no it's like constantly so it gets very difficult to debug so if you ever write a piece of code you know debugging is part of the business but debugging these things is very very complicated because you have three things in parallel the work may be boring the workers may be spammers and your task is ill-design okay so how we identify what are the problems so I'm going to show you a bit of a framework that at Microsoft we wrote a little paper little tech report and we have it in production actually kind of works so the idea is you establish a base signal so you just get everything are you going to pick the first thing I'm going to say how is my data set is the stuff that I'm getting actually so far useful people like it how it is okay then we're gonna check that the second we're going to do is like how are my workers are they spammers how's the quality of their work what do they do what they miss what do they get and finally we're going to assess the task design maybe my design is incorrect maybe I screw up no matter you know everything everybody's trying to do the best they can but I just you know I'm not helping them so since the previous speaker Matt talked about Twitter I'm going to talk about Twitter so my team we use so we have access obviously to web data but in case you want to join us we have access to Twitter Facebook quora Foursquare okay the only place on the planet that you have access to Social and web okay so we're going to label tweets very very subjective task and the point here is a very practical application you're going to talk about machine learning if you haven't talked yet on the class so we're going to build the classifier that's the idea so given a tweet I want to know if the Tweet is interesting or not because I want to build an index on interesting tweets and then I want to discard the tweets that are not interesting so then my index size is going to be smaller right looks like straightforward problem um we found that it was very very difficult because we couldn't get a lot of agreements it was very low integrated agreement and we tested our internal platform Mechanical Turk craft flower and we just couldn't move it so we came up with this idea so we're going to use the same we're going to borrow the concept of recapture okay which they have the control term so I have to guess one of those terms and the other one is the control they know the answer so we built this concept called hidden human intelligence data driven inquiries basically means that inside the experiment I have a question and I already know the answer and we're going to ask two questions one is an algorithmic question the second is semantic when I say algorithmic is I'm not asking you to code anything but I can compute this answer before I upload the data a little bit of this one so it's a production example here's a tweet okay the only task is you have to tell me if the Tweet is interesting or not that's our punch question number three but if you can see this the first question is how many hashtag words so words that begin with the hash are in this tweet okay so it's zero no hashtags one two three or more the second is a semantic question does the Tweet name a specific person so we're going to see if the audience can pass so for question number one how many hashtags we got on tweet very good pass for the second does it tweet name a specific person yes Paul Allen perfect okay so assuming you pass one and you pass two here's the question that we really want do you think the Tweet is interesting to a broader audience yes or no for doing that you know very very detailed systematic approach we measure the integrated agreement on question one pretty high so the kappas are between minus one and one one means perfect agreement zero is equal to flipping a coin negatives like this is so bad okay so for the first one Kappa 0.91 Alpha 0.91 second smell of Victory right so like get in there it's like very very strong and then it's like come on come on give me number three bomb okay complete failure again the workers are not spammers the data was good but we screwed up on question three there's something wrong with our design and remember when I say labels for the machines and labels for the humans I'm asking do you think the tourist is interesting to a broader audience yes no because my classifier is going to say yes no so I would like to build a model that outputs CS no so hey that's what I want but that those are the labels for my machine not the labels for the humans so we're going to give it another shot and instead of you asking yes no we're gonna do a greater grade answer okay so you have to tell me if the Tweet is interesting or not but instead of that I'm going to ask you to say is worthless trivial it's funny makes me curious contains useful information it's important news okay these are like these are labels for the humans not labels for the machine and right now we're going to see if the statistics are improving or not a we know how to pass one we know how to pass through on three although we didn't get everything that we wanted we can see that there's a bit of signal on the important news so people can identify the tweet it's about important news while this is not perfect then we can see where the problems were we can get a bit of signal I can you know produce more more labels and more categories and and get the numbers better but once we reach this then we're going to start building our classifier okay so once we get here after all this effort no we're starting to get high quality labels once we get high quality labels this data can be used for the rankers for the machine learning models for the evaluations for the constructions of the training sets right if you want to scale and repeat this is the way to go okay so all this effort in like 45 minutes just to tell you that everything is very hard and once you reach some sort of like level quality then you can do the other part of the machine learning so this is God give you data God won't give you the data which is grab the data we'll do the best we can with high quality labels and then we move into the machine learning part okay um the last part of the talk I'm gonna um move more into the research problems and some of the topics that you know the community is working on by the way you can announce you can we don't have time for uh questions the first one is the notion of algorithms using people so so far all the algorithms are like with machines right so I want to sort the list I want to find you know my nearest neighborhood in The Social Network I want to see the distance I want to compute uh you know a number I want to get my H index all those are like you put the graph in a table or some sort of data structures you run and I go and here's your number so we're going to do the same similar things but with people so there's a notion of Bandit problems in statistics which um a few days ago to a casino and they're going to play one of those bandage you know it's gonna you're gonna put a coin and you're going to move the arm and hopefully you're going to hit the jackpot and then you have so many machines so you have to kind of explore and exploit which one of them so the idea is to explore and if one is doing more as well just want to exploit that that Bandit so the idea is explore exploit paradigms so people are starting to work on also how to explore and exploit different set of crowds based on the knowledge and the idea is to you know optimize amount the amount of work by workers so if I know that I have a pool of workers who are very good at this German spelling I'm just going to ask you more German spelling if the guys are not good at German spelling but they're very good at Image level I'm just going to ask more image labels so how can I get that expertise that kind of knowledge in the crowd um so and this is important because humans you know have limited throughput so they cannot run 24x7 like a machine so it's very difficult to scale and you may want to make sure that if I'm going to ask someone to do something for me the task is very clear it's very clean description the payment is good so then none of us are going to waste some time that's on the algorithmic side and also when to stop when to stop when do I stop asking for something because it's I'm done or the task is so difficult it's no point the second is the notion of humans in the loop which is the part that I'm really really psyched about it because scenario that I work which is you know you have machines which are CPUs imagine the humans as hpus human processing units and the idea is to combine both together okay so I want to do a task maybe 99 of the task is done by a machine but there's one percent done by a human as a user I do not care I just won't get my stuff done so it's a little bit like Active Learning you're going to cover this topic later in the class but it has a double goal so the humans can check if the machine is doing the right stuff but also the machine can check if you are doing the right stuff as well and we use this pattern a lot for building classifiers on social data which are very very subjective so the machine does the best it can and then there's a human Loop that says Nope go this way or not go that way and with the data that the machine produced then it will check if the human is paying attention or not the last topic I want to talk about to you guys is the notion of routing this is a project that an intro bus did last year bestanushi from eth um so you can say I collect data and that's super cool but besides collecting data and just paying people what else can I do so the idea here is we're going to detect expertise on a social network so we're going to detect expertise from Twitter who are the experts on a topic on Twitter and then we're going to build something that we call a social load balancer so if you if you if you know a bit of routing algorithms on machines this notion of balancing you don't want to you know load one machine because the other machines can do stuff for you and that's great you know the machine will run 24x7 always giving you an answer but humans are not like that so if if you ask me hey crowdsourcing I'm gonna answer you but if you every single minute you ask me the same question or I get the same question I'm gonna get annoying and I'm going to answer this anymore so we built this kind of social load balancer that we detect a pool of people who are experts and then I'm asking one person answer me a question the next time I have to ask the same question I'm going to ask to another person and so forth which are expert on the topic so we call the system crowdstar and with a there's a reprint on archive if you want to get a paper but the idea is the following you have two tasks so have a task a and I'm going to Route this task to a specific crowd for example I know that given the definition of the task this is probably suitable for Twitter so I'm going to Route the question to the Twitter crowd or I may say this looks like a core type of crowd so I'm going to Route it to quora so behind the covers is an algorithm that detects expertise on topics and based on the query on the topic is going to say this this summary discount summary is going to say go this way or go that way and then inside each of this crowd is going to route to specific people okay so the question is I have a question on node Excel there's like is that for Twitter or for quora or Facebook or what have you say that the system says go to Twitter and then within the crowd who's going to be the first one um I did say or Matt is the other Professor right he's in South Africa it's not available so who is number third and so forth because you want the answer now and actually works so we we tested in quora and on Twitter I don't know if you can read this but the first question so we we created um an account on Twitter and we bootstrap it then we created the account and then we launch it and the question is can you recommend uh can you recommend us you know hiking uh parks for summer and you know we ask and there's someone who answers and you know makes a recommendation that's all good but we also learned that if you want to ask and you want to Route questions you have to be precise so you cannot say which are the top 10 bands what are the top 10 rock and roll bands people are not going to answer that question but if you say hey should I buy you know Led Zeppelin CD or Taylor Swift then you're gonna get an answer no so you know be precise in asking the questions to recap and apologize if I was going too fast crowdsourcing at scale it works but requires a solid framework fasten around easy to experiment few dollars to test you have to design the experiment is very carefully I cannot emphasize this enough it's so so important learn or try to learn everything about usability if you can lots of opportunities to improve current platforms if you want to work on getting better Emperor getting a better cloudflower regardless of the task you have to pay attention to three things workers work and task that's the little framework that I show you because at any time one of these guys is not going to work let me explain a little bit why so this is not like physics that every day if I drop the ball I expect the low gravity today I've asked something in them Turk I get an answer tomorrow I ask the same thing I may not get the same answer okay very very important history doesn't repeat itself in crowdsourcing okay being a mind of that and labeling social data if you work with social data is very very difficult difficult now students important to know your limitations and be ready to collaborate it's kind of impossible to know all these things in Industry this is a team effort it's not a single person running this so you have to have some knowledge of Social and Behavioral Science okay how how to ask questions uh cognitive load um all those things that you're going to get in different classes it's super super important here the second is human factors the best way to present the information you know color font text images third one is the algorithms how we're going to compute into reader agreement how we're going to manage the crowd kind of without the bad performers Etc then economics incentives how much you're going to pay we're going to pay for money you're going to pay for credit I'm going to pay it for like you know I'll give you a cell phone if you do this for me I'm going to incentivize people to work for me distributed systems you know how can I think I'm a micro asset is to really pull off people doing work for me last but not least one of the most important part statistics um okay short and sweet as my Twitter account and that's it so I'm happy to go through any of these topics in like depth and I can talk to you for hours I'm pretty sure you guys are tired of listening to me so uh whatever you want to ask me I'm gonna be here so just fired well yes sir sorry what was the alpha measure what was that correct vendors okay and then you were preparing that to K no that's a that's a um flight and hold on to be more precise so coins and flies are basically kappas that's okay this is Alpha pick any um Gryffindor Cloud scrippendorf wrote a book called content analysis so it's all here um coin and flies are papers from education and medicine because the idea is if you go to a doctor and doctor said you broke your knee and you expect the other doctor to say you broke your knee so then you're going to get agreement if people disagree then you know there's something wrong as well as educational tests okay so you get your you know sat there is it's very important to measure agreement more questions boring topic guys no yep so what's uh like the users are currently Big Data companies yes Microsoft yes Google LinkedIn what is the potential for wider Market do you see kind of a sustain a niche for data cleaning for Big Data companies or is there any kind of bigger new skills so here's my take and by the way this this answer is not sponsored by Facebook but Facebook for me is the crowdsourcing platform because every time you like you're basically telling this is good every time you comment or something every time you select an image it's basically it's they don't have a crawler okay so you guys are just recording the web for them and you're putting every little thing and they also have this very cool thing which I remember I was trying to I went to www conference in Rio a couple years ago last year and I was trying to log in from Rio and Facebook says Hey usually lock from the Bay Area not from Rio to make sure we know you they run a number of captures and the captchas were basically my friends and I have to tag them with data that was somewhat not ready or a little bit incomplete okay so it was a task that they want to produce new data and I was the worker I didn't I didn't get any I didn't get paid I just got access to the system so there's a lot of potential for that kind of work there's a lot of potential for creating new data so you may hear of the Google Knowledge Graph Microsoft we have also our graphs um how to curate all these things how to create new things new entities new relationships is a lot of that stuff is powered by human computation at some level in case I wasn't very clear on the example of the search engine when I say you are a computer every time you issue a query on Google or Bing and you click on position number one two or three you are voting if that page is good or not Okay so the way is the way a search engine builds the index is because you guys create a home page and then you put a little bit of a hyperlink to another page okay if you harvest the entire web then Diego can say oh this is good this is good this is good then you know page Rank and you have your order but at the low level those human annotations so you I created a link to Alex's page the machine cannot do that I did it so then they just Harvest on top of that um LinkedIn is is doing this obviously um if you're in LinkedIn and you have to um assess expertise of your fellow connections that's also human computation right so hey you're connected to Alexi you know is he an expert on data science yes no can you recommend this person for this job this is all human based I mentioned obviously Microsoft we do a lot of that Google does a lot of that as well as Facebook Twitter you guys used to it right yes you know what is a training topic right training topics the last bit is crowdsourced okay you have the Argos that will select a lot of potential training topics and then the pruning of the last 10 has a human Loop why not because to make sure that the Argo is good because I'm sure it's pretty pretty good but you're going to have a lot of things that it's adult content is not only because it's all it's obscene but it may be politically incorrect or maybe it's offensive so you're always always going to have a human alone obviously I'm just giving the presentation from the perspective of um what is called supervised machine learning where you always need something so I'm not talking about here about the unsupervised kind of deep learning which is like there's no humans in the loop it's all uh machine based but I think the the hybrid human machine is the way to go um so going back to your answer this potential opportunities right now I think we're just scratching the surface mainly data cleaning any more questions