Devreal

The practice of acquiring good labels

Event: Data by the Bay

data.bythebay.io: Omar Alonso, The practice of acquiring good labels

Recording: data.bythebay.io: Omar Alonso, The practice of acquiring good labels

alright thanks for coming to my talk I'm wearing the other shirt from text by the way from last year just to show respect so the usual disclaimer you know these are not Microsoft opinions on label quality just my opinions its joint work with a bunch of current and former Microsoft colleagues a quick outline I'm going to be talking about the problem of acquiring good labels in the context of big data and then I'm going to be talking about three main things the first one is wet web programming or how to program people how to debug and test a human computation tasks algorithms for quality control and then introductions sorry conclusions so I'm assuming you have a bit of a background on how to run human computation task and by the way for there's a lot of technical content so I try to add a footnote as a paper or a book so you can go and read those things in detail I try to abstract more important concepts for the presentation to minimize math and I wouldn't work okay so I'm assuming you have an idea of what human computation is which is basically how to use humans or people for things that machines are not very good at and this idea has been study particular by luis von ahn from CMU who is doing capture and also this notion of games with a purpose so you can play a game but under the covers were doing it is basically acquiring data for a bunch of our purposes you're probably familiar with platforms like Mechanical Turk or crab flower where you can choose create a hit which is a human intelligence task and we know you pay a few cents and people will do stuff for you and then you collect the data and without the data you to strain your models so in our case in this for this presentation I'm going to every time I talked about him is going to be about a worker you're producing some work or your radar or even Anna Taylor or judge but the idea is humans are going to be doing stuff for us a bit of a context it's about supervised or semi-supervised learning so this is not unsupervised learning supervisor semi-supervised we're going to ask experts or users to do the task the task is usually a little bit different from the real work because you're asking assuming you're going to do this so it's a lil bit artificial and there are too many important things from industry here one is scalability we're not interested in just getting in a hundred labels or a thousand levels for a paper this is massive scale thousands of millions of labels and the second discontinuity is not a one-off job you do this every day or every week or every month ok so the crux of the problem why we need labels labels are super important for information retrieval you need collections to start boost wrapping a ranker so you can search stuff on the web for natural language processing you need to tag all these things to get sentiment to get named entity extraction etc for machine learning obviously active learning artificial intelligence in particular if you're trying to build models there are a hybrid of humans and machines and a sample of common tasks at least today for gathering labels are content moderation the usual stuff in others adult content you have to build a classifier that can detect this information extraction please extract the name of the person from a page search relevant or the search result relevant entity resolution is IBM the same as international business machine and a bunch of other text processing this is a great book crowd source data management that was published recently which you know they didn't I didn't very good job at just doing running this survey just to get you know perspective from industry on how these things are going on so what is the label so label basically means that we want to get or tag a particular object with some sort of scale so if I'm looking for example for the for a page and the queries text by the bay and I get you know this page about text by the way is a page relevant and probably have three answers it is relevant very relevant somewhat relevant are not relevant I will consider that may be labeled or you usually have this typical spam email message and you want to know if the message is spam or not again it's a classification problem and the levels are yes or no however in the area of big data machine learning we have to be very very careful with labels right case labels are used to compute features that we use the features to build a predictive model and then we do optimization but label and experimentation unfortunately I perceive us extremely boring side of the house because everybody wants to play with support vector machines your networks you name it worked back everything is super cool but nobody wants to actually get that data so there's a tendency in the real world to rush the living process and I'm going to argue against death rushing I'm going to tell you why this is important so I took this picture from Twitter and this is a typical nothing against the datura camera this is a typical example that you'll find today in machine learning everybody is doing the engineering modeling we know visualization and how to do that but if you can see from the picture the gate data etl is gray nobody wants to work on that the talk is about how to work on that and in case you say well but you know who cares we don't need humans you know can do you know deep learning another thing you will always need a human either humans to do the training you need the humans to do the experimentation you need to view image to the devaluation it's a great article by the way a couple years ago and wire on how Facebook is using a third-party company to get all this adult content you probably heard from Ricardo's talk all the buyers are on the web and you probably heard the news from facebook on the on the bias on on the news and you can read the reading guides for trending topics etc etc why it's important because there is a life cycle of the label i'm going to use information retrieval example here so an information retrieval usually you need to ask people to annotate a collection of relevant relevant content and based on that you build a rancor and so forth so how this works we start with the following we have a data set so we have a collection of documents and i'm a judge and i have characteristics so the documents already have some characteristics maybe the documents our politics or sport as a judge I also have characteristic maybe I do like support maybe I don't like politics my task is to do judgments I'm going to do judging on a document and a query and I'm going to produce a judgment because I'm a human and I can have an error another person will do the same we're going to do an aggregation of those judgments I'm going to produce a label okay that part is basically using a crowd to label a data set the second part is we're going to use the label to produce some learning model we're going to learn the model so we can predict the rest of the labels and with the rest of the labels you can annotate the rest of the collection also the labels are going to be used later to evaluate the performance of my system so what I'm trying to say in the picture here is like they're everywhere we got garbage in garbage out is the same problem if we don't get the correct levels early on then there's a lot of implications down the pipeline so label quality okay what do you mean by quality meet internal customer needs and it's free from the efficiencies that's our definition in our case you want to process where we want rush labeling or Nagano outsource you own it end-to-end you can crowdsource so you can outsource the acquisition by humans but you own it the entire process its large scales and it's continuous and it's another dimension to even make things more complicated there's a spectrum of labeling task okay so we can start with something very very simple which is what we call an objective question that has a correct answer then an aggregation approach will be to you know make sure that reliable judges assign an appropriate level for an item and then devolution technique is usually done against a ground truth or a goal set that's great we come we complicate things little bit more because you're going to have other cases where the task has the best answer not a unique as before the best answer is partially objective and they're going to use interval agreement to actually measures and determine which is the label for the item and then we evaluate the technique basically a billion in the workers by comparing the individual results with the consensus of the rest of the crowd and there are other cases were this is very very subjective so you have to use something that we call repeated Horning to determine the probability of the label of a label of the idea for an item and then what we want is to evaluate the workers by computing the consistency of results between groups so those are the three type of labeling task you're always going to encounter and we're going to use a programming approach for solving this problem so plumbing is hard and there's a lot of very good books in particular this book by brian kernighan a rope hike on the practice of farming i'm going to take a similar approach for writing code that is executed by machine some people this is the premium view for humans and machines okay we call that I call that wet performing and there are two two important things here is you know Don Knuth set we know machines have no common sense they do exactly as they are told no more and no less however is we are humans and we you know we do a lot of errors so how can we hit deal with these two things so we compare humans and machines humans executing code you know we're not that great so the instructions said for humans is somewhat unknown we take our time so there's latency we need a cause or incentive for doing something we obviously produce errors task difficulty is an issue with humans and their assuming human factors are maybe get bored I'm distractor I'm Lacey etc so how do we start the first one is asking questions the same way you instruct the right code this is part art and part science the instructions are key workers may not be expert so you cannot assume that they have some understanding in terms of the terminology maybe they don't understand relevance we didn't understand what you mean by label so they understand that you always have the show examples and if you can hire a tech writer or someone who has an English major actually do this for you and then the you know the engineer will will write a specification by the actual tech writer will write the hit for you in terms of the heat design the experiments should be self-contained always in the same place keep it short simple to the point clear with the task engage with the worker all your tricks and things that you learn from user interface design you know document presentation and design super crucial you need to grab attention so have good content and always look for feedback if you are doing multilingual task do it in English first and then you scale to other other languages other the science principles very quickly in your alignment legibility reading level this is very very important people are going to read this so how difficult is to read to parse text there are multicultural multilingual issues as i said before remember you know color blindness the cognitive load there certain thing that takes forever to actually understand an exposure effect which means that don't ask them to do always the same thing because we do the same stuff over and over again we tend to like it how we measure this once humans are performing the task as i said before we have to use redundancy so usually we have to ask why more people to provide the answers for us so then we can measure what we call agreement reliability and validity the first one is to measure the inter agreement level between the workers so you can do between the judges or the workers or between the workers or judges and the gold set statistics for this usually coins kappa for two Raiders flies if you use in n number of Raiders krippendorf for any number of Raiders and in case you have missing values you can use this flies and coins doesn't work with missing values and also you might need to use an internal consistency stat this one very popular called care 22 the Richardson as well I'm not suggesting that you have to use all of them but the point here is use one that fits your domain and needs you have to measure you know how the workers are doing against your data said you have to measure with something this is an example I'll have to measure and by the way to not use percentage agreement can never use that always flies kappa or krippendorf KR 20 other practical tips sign as a worker and do some work you have to eat your own dog food try to address any feedback as soon as you can poor guidelines payment passing grades etc everything actually does count now there's a human side here and the human studies as a worker and myself if I have to do some human computation I do hate when the instructors are not clear and this is a lot of cases people will say oh I know how to do my hits very well incorrect they have no clue and this is always the case people don't know how to write good hits sometimes you're not just a spammer but you actually don't get you know what the requester is asking you to do oh maybe the task is boring and in terms of incentives you always want a good payment but sometimes there's something else maybe you're learning a new trick the content is good you'd like to you know look at images and that's what I'm doing this image tagging for you as a requester your problem is situation you always want to have a good crowd that does the label for for you so you always want to maintain the same force and it's a delicate act because you want to produce the right results but at the same time you want to you know have good content that is appealing to workers and obviously you want honors answers for the task you don't want fake answers and you want qualified workers but sometimes you don't have them and managing crowds is very very difficult and it's a daily job okay it's much harder to manage the data center there are number of design patterns I just outlined some of them probably the most interesting is called the fine fix verify that's been designed by Michael birthday who is now that stand for and the idea is you ask the crowd to do something so find this fix that thing and then you have another person who's even very verify the fix it's kind of a workflow and works very very well now we can always go further and say this is great but again from the priming perspective let's let's kind of think in terms of program structure and design hits that humans can do very well but in terms of data pipelines and workflow instead of having one single hit for everything and then go to the machine learning model just do little pieces at the end you get the label but all of them are like in your little little things that you do as you go to get your final level and this involves combining humans and machines this is a great example a project called cascade for taxonomy creation using crowdsourcing and the beauty of this project is basically the provide and design three main hit primitives so the same way as you have primitives in a prominent language you have three main tasks to generate select the best and to categorize and they all work in some sort of a workflow one of them on the hits and then at the end they have this algorithm called global structure inferences that is going to you know get all these data gathered by humans to produce a taxonomy and the taxonomy that is that has been produced is very good to el día in terms of quality and price so you can get more details from the paper here's a project that we we we that Microsoft a few years ago on near duplicate detection so near the placate detection is one of the most important promise on the web because you want to detect if to web pages are not only exact duplicates but if you are near dupes you know there was a very good algorithm and we wanted to know how can we perform the evaluation on this near dubai go so we do the following we use two platforms to generate labels one was you at rest which is an internal microsoft platform and then mechanical turk and we use a two-phase approach first we thought okay come on this is so easy you know it just give you two documents and you have to tell me they're similar or not it turned out to be a very difficult task because people don't agree what a news article is is and what do you mean by near dupes so what we did is we we did two two phases the first one is we do kind of in a cheap way in mechanical turk we go through all this document and ask the judges is this document you think the document is news or not yes no that's it and from there we use a separate platform you a choice which more qualified workers to actually assess if the near duplicates the tube sorry the two news documents were near duplicates or not this allows us to do the following separate workforce so one workforce we pay less and turk because which is one binary yes no news article the more expensive workforce in your hrs to actually do the near duplicate evaluation and this allows you to paralyze so the same time you're just getting all your news levels if you get candidates you can already start evaluating the to do the near-duplicate detection and you know all the details the paper by the way the data is available to download and also the labels on all the heat templates you can replicate all the study ok this is just for writing instructions this is going to get a little more complicated is how to test and debug I mentioned and I was quitting turning an umpire on plumbing and there's a great book by fred brooks juniors on the musical month month and you have to always throw a prototype and unfortunately human computation we don't do that sorry so what's the problem in testing we always attempt to break a program and in the beginning we know the problem is broken and which is not to find out why it's broken now if you come if we compare machine computation base versus human computation at the design level when you're using machines you know you're going to throw away the prototype you do systematic testing and if something is wrong is your fault okay you don't say it's GCC is the library is my pc or mac sure whatever it's your fault human computation at the sign level we are returning to throw away the product type because a we're going to which is paid twenty bucks in am turk and you know why why we have to throw this away the testing is ad hoc yeah we cannot test sometimes with tears sometimes we don't test and debugging is it's the workers problem you know there are spammers you know they didn't want to work they're lazy people well the new series is your problem you know this is the same as you know debugging code you have to do is systematic you're going to throw away a lot of hits and for debugging purposes first is your problem and then we'll see if it's the workers problem but before doing that you have to have a development process and the same way you're going to write code using agile or water for whatever it is your metaphor you have to use some sort of development process the development process that I'm just suggesting here is just a suggestion you start with a prototype which is a very small data set for your own internal team you do the testing on debugging you do all these measurements you do this interior agreement between yourself even and your two other colleagues and co-workers Eve everything works fine then you go what we call an early stage production which is still a small data set always small data set 101k items crowd base mainly for calibration do the workers can do the same as you can you measure all these agreements yes you're throwing away 100 bucks what's the big deal sorry and after that you're going to go in continuous production which is everything works fine you tested everything and then you actually go in large scale you're going to partition the data it's going to go crowd base fully crowd base and then you can reinforce quality control so how you test the hit you send you follow the same methodology as code I mean you have to test for validity and possibility if you had a problem with the head you don't patch it okay you rewrite it from scratch you're going to version all the time place and all the metadata because maybe this kid was for three people maybe 45 people will you pay a dollar maybe pay fifty cents all those things however things are going to get very complicated because debugging human computations task is usually very very difficult why because you have multiple contingent factors excuse them for the world we have three things so the workers maybe they're spammers maybe your data sucks nobody actually do any labeling and all maybe the task is actually very difficult to to perform any work on it so you have these three things in parallel and we came up with this framework that we're going to iron out each of these elements first going to play with the data they were going to play with the worker and they will play with the task so it's kind of a slight checklist to make sure that the data looks fine the worker performance looks fine on the task looks fine once that is done then you actually have debug and test your human computation and some playing us it was mentioned in the previous dog is super crucial here sampling from good people from gun content okay so how can we do this so we came up with these technical headings which is human intelligence task data-driven inquires which is basically we borrow the idea from capture so you have the control turn and then the actual term that you want people to to enter because it's the label you want so we're going to automatically generate questions from the data set one is an algorithmic question very easy to compute the answer and the second is more of a semantic question those are mostly to see if the workers are alive they're doing what they're supposed to do and they go for the third question which is the question that you really want the answer so here's an example on a very difficult problem which is to eat classification I'm saying very difficult because when people talk about labeling social data it's very very hard but it or not is very very hard so the tweet here is we want to know if we can derive some categories from a tweet very very simple so we show it tweet the branded no account just 140 characters and we have two questions those are the hidden questions the first one is how many hashtag words the words that began with hash are in the tweet so no hashtags one two three or more so you can compute the answer already so when you upload the data you know what what's the answer for the first question and it's just a kind of a sanity check to the workers the second is more of a semantic question does the tweet name any specific person in this case is Paul Island so the answer is yes so you can use a named entity extractor to see if this is correct and that the third question is the question that you really want the answers and you want to know if the twit is worthless trivial funny makes me curious contains useful information or it's important news and as you can see I'm using these soft labels so i'm not using yes no i'm not using a Likert scale because those things don't work you have to find labels that actually work for people now we spend a lot of time working on these questions because the question number three was a programmatic question question one or two are working very well but we started using yes no we used to like your scale and we couldn't get the labels that humans were able to work better until we found this labeling another example is to use the micro breaks so basically diversions as part of the task so you do some task and once in a while you have died versions and this is Dana Gould and the show that is very very useful and can improve the work quality even if you sorry in the case that you have Micro breaks that are a part of the context are even better I'm going to touch something that is also a part of the bugging but it's the other way around i was i was talking how to debug tasks that are part of humans and machines you can also do what we call human debugging so we're going to use humans to debug the ml model that you just built we call that beat the machine so you're going to find cases where the machine is wrong in particular what we call the what is called the unknowns unknowns things that the model believes that you know this not doing a lot of errors but actually the model is wrong so not only you can use priming techniques to identify bugs but you can also use humans to identify bugs in the modeling and since I have 10 more minutes I'm going to talk about migrants for quality control quality control C is an extremely important part of every human computation task it's not just the workers are bad but is bi-directional so you may think you know the work is our spam ours but the workers may think you're a very lousy requester because you don't know how to ask questions when yes s work quality before the task is uploaded during the task after the task okay always always be assessing quality so before that you can you run qualification task for you know screening selection of people during you can also have random checks to make sure that people are doing what they're supposed to do at the end you just run the number of metrics and you make sure that people have passed on the answers same as a test how do I measure work quality we can compare against another person another label where you can build predictive models of the labels and you can also try by yourself or you can use tears approach of different workers comparing two known answers this is called kind of gold set honey pots um you have an answer and it is what I make sure that I pass or not the problem is you have to produce these gold sets are expensive you have to have experts are going to produce the goal set and you have to maintain them because people may have already guessed the goal set so next time we do the work they're going to already know maybe do they pass the answers to other workers comparing to other workers that is called consensus redundant label or plurality it's a one no metric for measuring agreement again in terms of cost benefit is rendered nancy's you're adding stuff they already know to this qualification test just I'm going to keep briefly advantages it's very good but disadvantages you know it's requires a lot of work for you to build the qualification test and some people just don't like to ask a test before actually doing some work now if you're going to go through for the algorithms using in practice majority vote actually works very nicely there's a school of thought that you have to use emx Peck taoiseach maximization to do all this stuff I get another level is another girl algorithmic for getting the the end result for the labels for all these workers vox populi very good for cleaning data and then paramedic go this was done by the guy said crowd flower we have microsoft been working with a call adaptive algorithms and the idea here is to use explore exploit approaches with what is this is kind of a quality cost trade-off and the idea is there are certain cases that you don't need a lot of workers and there are certain cases that you need a lot of workers so the point is how many workers and want to stop so if the query that I have to evaluate is facebook on the website is facebook com maybe I need one worker but if the queries solar storms and this is the website and I have no clue maybe I need five so when to adapt more or less and the adaptation is going to be based on the complexity of the hit some of those thoughts going to be very complex so I need more some of them are very easy need less so the idea is to define stopping rules which is basically algorithms that no one to stop this asking more less labels and I want to stop if I have historical data for all these workers or if I don't have historical data and if you can do this you can build ground truth automatically full paper is going to appear in cigar in a couple of months now there are some practical considerations here so all the expectation-maximization approaches they are very good in theory but in practice you have to have historical data and sometimes this is hard and you can have this kind of coaster problem because you don't have historical data aggregation rules based on majority vote is actually a very nice alternative you can use you can do majority you can dish you can use unanimity and you can use quota so majority if you have five workers and I have three say yes to say good its majority but you know it's not enough so maybe you want to use Korra so you want seventy percent of the workers to say yes or seventy percent of the workers to say no and finally you can do something completely different which is behavioral features so the idea here is to fingerprint the task so you're going to trade the worker using javascript to know where they put the mouse whether it click when they stop when they took less time more time and you can do that getting the metrics this this notion of fingerprint is going to tell you the fingerprint of the task and you can compare others workers to see they're more less in the same finger finger print or not fresh out of the oven there's paper by the Google guys on a system called Wernick they're actually use it's a very nice tool for annotating information extraction they're using a weighted majority vote all fingerprinting and they show that you can do better than any of the am based approach if you just use behavioral data so far this is great but looks like a ton of work poor amine is hard priming for machines is hard plumbing for machines and human is harder so this should not be unexpected remember we want good labels and data quality and experimentation of designs are preconditions to make sure we get the right stuff you know all the labels will be used for rank your models evaluations etc don't cut any corners I show you the Brahmin side so how to code a hit you can use patterns you can modernize same we're going to chop your piece of code you can do all these pipelines little workflows you can you have to test on the back you have to maintain and you have to monitor the work performance last two slides the first one is if you're going to work on computations that involve machines and humans in synchronization it's a it's a very delicate balance but there's a lot of potential now the most important part for you guys as the designers or architects or data scientist you have to know when to use a machine or a human for computation labels for the machine not be labels for the humans the usual approach is I have these features that I want to build a classifier and people just take that and make it a web form in mechanical third to get the labels those are labels for the machines may not be labels for the humans find those labels then debugging code executed by humans is hard but you can do it at the same time why don't you choose debug the machine model with humans so it's a win-win you're going to debug humans but the same time the humans will help you debug your model and the final bit is the best arguments for the machine may not be the best choices for humans ok so there's heavy data-driven are girls that worked very well but sometimes might not work with humans so take aways and I'm just done with this reputable label quality at scale works but requires a solid framework use all premium principles always use premium principles if you can data management three aspects that actually really need your attention workers work and tasks design and you cannot do this by yourself ok there's a lot of skills and expertise required social the hero science human factors algorithms economics digital system stats that's my email my twitter handle thank you very much for the time gracias por la presentazione Amaro amar allenton since the talk was a little bit top down I'm going to sum it up for you so you either memorized the whole talk if you want to acquire good labels or the other option is hire a German psychologist they do this stuff for a living okay thank you very much so for Q&A try to keep your questions short because we are trying to set up everything for the next speaker and I'll come to you with a microphone yes first of all I would like to thank you for the talk I've been working in the area of active learning and all you said was very very relevant to what I'm I work on and then the biggest question is with given the number of people that work on social data and as you pointed out getting quality data quality labels on social data is very very tough why isn't there more focus on the kind of work that you talked about today like getting good quality labels why is there more focus on the type of machine learning models were rather than the labels thank you very much for the question that would have been my question too great question and thanks for asking the question i think the the answer is everybody wants to work on the modeling which is kind of the sexy part you know features you roll these metrics create stuff but as we acquire more data and we're going to find we're starting to find out the human sanitation all the lebanese getting more and more important facebook example is one of them okay so even they leak the editorial guidelines you can read them is around 30 pages and people are questioning that you don't get the right labels diagrams are not going to work perfect as Ricardo dimensions before right so you can have all these biases and you have to correct and a lot of cases you need humans in the loop there's I mean this is going to get the worst it's not going to get easier it's going to get worse why because we have Facebook Twitter data who knows what we're going to have next week and next year and those things just get what I guess what they're going to get even harder because everybody use the same thing Kimmy data features labeling optimization and prediction and everyone knows how to move from features upwards on the data well someone will do it but you have to own it if you don't own it you won't get good quality in my opinion so this is a thing it's just getting started thank you very much if you want to tweet about this awesome