Devreal

AI at CrowdFlower

Event: Bay Area AI: CrowdFlower AI and Text Generation

Bay Area AI: AI at CrowdFlower, Lukas Biewald

Recording: Bay Area AI: AI at CrowdFlower, Lukas Biewald

so thanks so much for coming out on a what is it Wednesday night really appreciate it I always kind of have trouble preparing for this stuff because it's so hard to know the kind of technical level of the audience I'm just curious how many of you guys are would you say your engineers got it so most of you and like how many of you use machine learning and you're in your daily job how many how many of you like built a machine learning model somehow okay cool thanks that's that's super helpful so I'll say my my background my background is in machine learning and I I spent you know the the first part of my career doing lots of things to make algorithms better but then I started CrowdFlower which is a training data and human-in-the-loop platform for machine learning and I wanted to talk about why I think that human-in-the-loop and training that is actually so important for making machine learning work I was talking earlier to some people about you know what would I work on if I was trying to get into machine learning you know like what algorithms would I would I suggest that somebody learns and I think it's I don't know if you guys would agree this but I was thinking it's kind of ironic but I would spend maybe the least amount of my time on the algorithms of machine learning right I think that's the part that's actually kind of the most well understood it's where the tools are the best and it's everything outside of that where it's actually really hard and we're actually I think people who have really worked on machine learning for years have a lot of skill but then people that are getting started really really need to learn this so I'll show you I'll show you exactly what I mean and I think one of things that Ellis traits are the best is this cattle competition that we ran a few years ago and if you guys done a cogwheel competition yeah which one yeah which one Oh malware costs what nice yeah yeah why prices nice Kaggle yeah so so catalysis website where where you can go and they they basically companies will submit datasets and they'll say whoever can model this the best wins a prize so it's really awesome in that you get lots and lots of people around the world working on your data set and trying to model it the best so you know at CrowdFlower we have tons of data sets and we thought it'd be fun to submit our own data set to a competition so he took it we took one that our customers care a lot about this is actually search relevance for e-commerce so saying if I went to a shopping site and I put in a search query how good would my result be and this is actually the stuff that I worked on in industry before I was running CrowdFlower and so I was really interested in how good the results would be so I mean Kaggle is amazing crowd sourcing is amazing you put the stuff up and you immediately get people working on it right so they had a metric of accuracy and this is basically predicting how relevant a search result is for a search query and so if you just guessed you get a little bit above 30% right actually if you guess the most common a category you'd get a little above 30% right so you could be worse if you randomly guessed right but a simple baseline is is just guessing the most common category and and so I submit that in the 11th of May and then on you know the 12th of May within a couple hours you see that someone's already doing a little bit better than guessing right so you know so it was clearly done something with the data set and this feels really good and then in the next the next day someone got up to 50% accuracy right so you know you start to see really quick results and then 14th of May 15th of May this gets up to 60% a little above and I was thinking you know I wonder how good this could be you know cuz the inter annotator agreement on this was about 95 percent so you know theoretically if you had an AI that was as good as humans you get up to 95 percent and I was thinking you know maybe they'll get to 80 maybe I'll get the 90% I wonder how well people do because we had tons and tons of people around the world submitting these these datasets but actually what happened was this right so you know after the first four days in terms of raw accuracy there was not a lot of progress right so you know the best result on the 18th of May right that's about a weekend was under 70% and it never got above 70% and so in terms of raw accuracy this didn't this didn't improve very much but actually in terms of sophistication of the models it exploded right so this is the number of participating teams over the lifetime of the competition and you can see that it went up and up and up and actually this doesn't even include the fact that each team was submitting tons of iterations on their model and actually I was in the message board so I saw all this right people are getting really detailed about how we collected the data they were going into more and more sophisticated systems more and more feature selection kind of doing all the things that you'd want to do if you wanted to label data more and more accurately build better and better models and I was actually amazed I think our prize money on this was about $20,000 I thought it was incredible of how hard people worked for this prize right so I think we really see state of the art in machine learning modeling this data set as of a year ago is about 70% accuracy but you know if you think about this and I think that we've all seen this curve I'll show you one I'll show you one more kind of famous one from the do you remember that anybody know the Netflix price hey people remember that so Netflix prize is kind of the first coggle competition in a way so Netflix offered a million dollars for the person that could improve their recommendation algorithm the most and so they actually ran this competition over a couple of years and you see almost the same exact curve right because it's percent improvement you know it it basically it basically flattened out right away right and actually there are these points that we didn't have where you can see that there was a breakthrough right so if you know people were working on it for months and months and months there's no improvement and these are thousands of teams submitting their their results and then finally one team would get a little bit better right it would be like a huge deal but in the end the researchers doing this were like working full-time on trying to win this competition and you know I think I think at the time that the Netflix was out I was working in machine learning I was thinking you know this is so exciting it's like so fun to try to make these these algorithms but I'll say managing people that have returns like this is like maddening right because you know weekend you say okay I've gone from 40 percent accuracy to 70% accuracy what does that tell me right if I'm if I'm writing code I at least have like some sense they're putting more work in will kind of take me closer and closer it's a bug-free or like you know ready to ship right but with these machine learning projects you know you might work another year after going from 40 percent to 70 percent and you're still at 70 percent accuracy right in fact we see with with the competition we ran that could that was probably like a hundred or a thousand man years of work that were put in with with basically like zero returns right so I think this is one of those things where people find it really difficult to to invest in in in projects and actually you see things where you kind of understand why AI goes through these hype cycles right where there's like all its enthusiasm when a new project starts and then the enthusiasm dies off right you get these AI winners mmm which way may be headed for again right like this is the there's actually the Google car so they have this great metric where it's miles per disengage right so this is um how far the car can go before a human driver had to take over and actually grab the wheel of the car to keep it from crashing right and so you can see that that if you look at the the four quarters of 2015 they made incredible progress right it actually this isn't flat for a long time and you see that that they're basically doubling the miles between disengage each quarter which is I think like this is a harbinger for why self-driving cars have gotten so incredibly exciting and popular right because you know things really starting to work we're kind of one of those moments in that in that Netflix contest where people are kind of making breakthroughs right and and you know learnings working really well in these vision applications but if you actually look I checked how far a human goes between fatal collisions and it's almost a million miles right so you know each disengage isn't necessarily a crash it's not necessarily a fatal crash but this is still pretty far from a fatal from from basically human level performance right and actually this is driving in Mountain View right which is like I think you almost couldn't pick like a nicer place to drive a car right like people are super polite I was actually just in India and like I can't even like I think if you put like one of these self-driving cars in the street like people would like immediately like destroy it right like you know so it's like it's actually a very contained environment like it's always sunny so so I think like even even in this environment you see like interesting progress but it's really far from being a safe thing to deploy right but I think like what what Tesla and others have done is they said okay we don't have a thing that that we can use right now and you may be really far from the thing that we can use but we can say okay in certain situations we have really high confidence that it's safe enough we can actually make a shippable product right and so you know you see people call adaptive cruise control I think Tesla kind of took adaptive cruise control and boldly called it a self-driving car but but you know highway driving actually works super well right and if you have a good system if you know where it's safe to go and you have a good system that alerts a human being to take over the wheel you can actually take a car that's maybe only like 50 percent accurate right maybe only if it can only drive in the highway it may only work in like you know 50 percent of the situation's that humans actually drive cars but you can actually make a useful safe system out of it not through any algorithm provements but through human-computer interaction right it's that's really what I want to talk about because I think that's like so powerful when you notice my favorite example of this is advanced chess which is something I learned about a while ago I'm like a huge game player and Kasparov beating deep-blue is like moment like in my life like I was so excited but what's actually amazing right if computers got better than humans at chess when they beat Kasparov I mean it's been like 20 years ago now I remember watching the games is a long time ago but actually up until last year the best possible system for playing chess involves computers and a human operator right so so actually 20 years went by where humans still had something to add right so so these people would play this advanced chess where they would actually like look at a computer running multiple computer programs and they would take over in situations like kind of closed games like highly strategic situations that weren't tactical and they would actually use the things the situations where humans are better to play it better right and so I think that's it just shows you kind of the lag right I mean it felt like computers really blew by humans in terms of skill but even a domain like chess which we now kind of realize is designed for computers to work well humans still had a lot to add for a long time so I sort of human the loop paradigm exists for a lot longer situation than you'd expect right I mean probably for 20 years before that for most chess players having a computer involved would have helped you right so this kind of like a 40 year window where we can actually collaborate and I guess now chess programs have gotten so good and so effective that humans just like screw it up so so the best program now is like actually a chess program but it doesn't mean that um there might be ways to still work together better I don't think a lot of people are looking into this so you know I think like okay what's the what's the kind of design pattern here that we can take from this right because you know these cars this advanced chess it all kind of feels like very different very different systems and so you know or we got to like kind of come up with something custom for every different kind of application and I think there's some general there's a jacket general design pattern that you can pull from this which is today's okay simplest way we can do it we start with an AI classifier where it's confident we use the output where it's not confident we get human annotation and then we do active learning right so it turns out that the places where the algorithms not confident are actually much much higher value for training the algorithm right so the places where the humans have actually intervened they're by definition cases were the algorithm struggling which is actually we see and and there's papers that that corroborate this it's like 10 to 100 times better use of your money right so even if you were ignoring this output right and all you're doing is trying to make a better classifier this design pattern is still better because you get better human annotations by picking the cases where the algorithms not confident but you also get out of this is a system where at any accuracy the a classifier it's useful right I mean if you have an 80% accurate classifier if it knows the 80 percent where it's accurate that could mean 80 percent cost savings right because this case is if it's confident those 80 percent of the times where it's right you can just use it right and and all good machine learning algorithms have a way that you can infer a non-biased confidence value from them right I mean some of them have seems like a little more contortions but every good every good system you can get you can basically build this design pattern on top of it and so I think what this means is that you know for lots and lots and lots applications this is actually the way that things get deployed right because a crowd floor what we see is how companies actually take systems and actually launch machine learning systems and there's always a form of this in almost every machine learning system that we've seen actually get done so I think this is really something that machine learning people really need to think about right because I'll tell you I think that's probably obvious to the people in the room who work on machine learning but it's not obvious to the CEOs of companies that I talked to that this is even possible right they haven't thought about this it's like you know they they you know I think a lot of times people don't trust these confidence values because if you think about human beings like you should not trust human beings confidence values right I mean like you know there's like that famous psychology effect were like actually like more confidence is correlated with being like less accurate right so I think I think a lot of people who don't work in machine learning have an intuition that it's really dangerous to trust confidence values but you know advantage of of AI is there's no ego so the confidence values are actually useful for you and I'll just say like you know there's like I feel like a lot of papers that I come across you really have to kind of squint to see the improvements as far things like even something like deep learning which can be amazing right is often kind of subtle improvements but active learning right like picking the examples where the algorithm is not confident is something that's just the the effect is so massive right I mean it's like it's like 10 or 100 times more efficient use of training data so you know for a lot of people if they trained it cost money to collect that's like 10 or 100 times more efficient use of your money right so I'm surprised that there's not more research around this and more thought around it for how powerful it is and I think if you look at kind of like some of the really smart successes you see a lot of active learning you know I was talking with the the alphago people right and and you know I think like one thing that's really cool about alphago compared to compared to the the deep was the chess program deep blue yeah deep blue I mean it's like it also goes such a modern approach right like it was trained on watching lots and lots of games of experts right it was actually trained on training data and and I think it's it's actually really interesting to watch the way that it failed right so you know deep blue I think it played very computer like game where's alphago so I'm uh I'm a passionate go player wanted to be professional go player when I was younger and so I think I can really say this deep alphago really played like a human being right like I think I played like a very very good human being but it really put like a human being and it was interesting too it made mistakes like a human being so you know game for that it lost it actually played like a shockingly bad move then I think like I think most humans actually like most expert humans would have been able to to catch it and tell it not to do that so I actually think it's very clear that with go even though alphago is clearly better than any human player I think that at least now it would benefit from a human Lou design pattern still right and you really just saw that in game forward it kind of went completely haywire I think there's another example this which is the the postal service I think this is like such a cool example cuz you know 1982 they actually launched OCR I didn't even I didn't realize this right but like they used those TR in 1982 and and it basically only worked if you like printed in exactly the font they wanted right but they're able to get like 10% of like you know the huge mailers to just like print like exactly how they told people to print and so they were able to OCR it and that was like awesome for them right because they saved like 10% of their costs by not having to read to let the letters and I was also I was kind of like looking into this like geeking out in it and a lot of people say that with deep learning OCR has gotten better than human beings and that may be true but the postal service still actually has two people that look at really badly written letters so the postal service still believes that they need to have here in the loop right so this is like a but but actually it's uh they publish the numbers it's over 99% of the letters that go through actually get read by a computer but there's still there's still exception handling with a person really no way oh man right yeah well I look at my grandmother's letters I'm well I'm like wondering like man just this is probably the one that got a person to only actually have 98% accurate over the stone so to two percent still get handled by a person and you know they sponsor some of these like handwriting things so they're not like messing around I mean they really want to make this work and I think it's like kind of funny like postal service launches essentially a machine learning thing in 1982 with human loop and if they had waited for accuracy to be higher than humans before launching an OCR system they would still be waiting right so it would it would be like over 30 years later 35 years later they still haven't launched it right so they were able to like ship this thing 35 years early by doing even loop and they've been like benefiting from it slowly as it as it collects more and more data and then like one thing that we just see a ton of is is chatbots cuz they totally benefit from this this paradigm also right it's so frustrating when a choppa doesn't understand what you're saying there's so many edge cases but there's also just so many easy examples right I mean you can make a rule based chatbot in a lot of cases that catches so much of what people say but I think to make like an actually human like chatbot you basically have to solve all of AI right so the gaps enormous but she can fill it with with human the loop right and so you know I mean not surprisingly CrowdFlower that I run is a human loop platform but and and so what crop flour does is it lets you do all kinds of different labeling right setting up jobs where humans can do the labeling and make sure that it's accurate right so I don't want to give you a sales pitch on CrowdFlower even though I just did I can't help myself because I just think it's like it's so important and it's it's it's often like the missing piece to making the difference between like AI kind of being like a nice idea that's something that people actually use and we see it over and over but I'll even show you like just like backing up I mean this is I'll show you that in graphs the reason that I started CrowdFlower which is this and and like you know I'm a math guy I worked you know my first research job was trying to make better algorithms and it's like incredibly it's incredibly frustrating experience right I mean like so I just I took some toy day to here on the left and I compared naive Bayes to maximum entropy to SVM for some some text classification and you go from like you know 90 percent error rate to like 17 percent error rate with SVM which is a little more machinery I mean these days I mean these these graphs for like years I need to add you know some kind of like recurrent neural network to this to like make it monitor bump telling you like that'll be like 16 percent like accuracy like you know I mean and and like you know you you get like one percent accuracy improvements and you you publish a paper I mean I have papers published with like less improvement then then you see here right and but then on the right you actually see kind of the same thing where it's like here's a paper where you know the redline is kind of random like a sort of stupid strategy in the green like a smarter strategy and you see it's like okay like you know there's there's a you know maybe 10% error reduction and that's like you know gets a paper published and so what you end up doing when you work in the real world probably what you guys spend most your time doing is feature selection right because feature selection actually really works right so you know here I go from you know grams of my data set to by group by grams to you know grams and by grams and you know each time it gets a pretty significant improvement and this is like I think this is what like you know people actually working on machine learning spend like 90 percent of their time doing right it's like you know how do we segment this tax like should we include all the punctuation I remember we had a customer that they they recently switched actually using the emojis and their data set and got this like massive improvement which is like totally made sense in hindsight you know but it's like you know if you're not really looking at the data you wouldn't realize that there's like lots of emojis and all of a sudden that you're encoding wrong but that's the kind of stuff that like you know you try like a hundred different strategies for for feature selection and like two of them get you that huge effect and in retrospect it seems obviously they worked but it's like I think no one has a good system for like a priori knowing what what features to use and you know deep learning has this promise of you know kind of doing the future looks and automatically but that's not what we hear from our customers who are actually really trying to play the stuff so far Sarnia so people spend all this time on this but there's actually like a thing that you could do that almost always leads to improvement that for some reason like feeling this mental block like they never think of it which is just collect more data write more training data and so you know this is the same data set I was I was that I was talking about in here it's about 10,000 so I went for like 10,000 to 20,000 to 40,000 records that I fed into it and each time it actually got more better than doing really good feature selection and much more better than going from like naive Bayes that kind of stupidest algorithm you can think of to SVM which is like a little more you know state of the art circuit like 2,000 algorithm and I think like this this I really like this graph to it because I it's this is from a paper that I think it's pretty funny where here the the x-axis is how much data they fed in and the y-axis is the accuracy so down is better right and so what they're trying to show here is that this green line is better than this red line right and that's like the point of the paper but they're like totally missing this incredibly like obvious effect that they're like accidentally demonstrating which is that like adding more data is causing their algorithms to get like way way better right so their point is that at each fixed amount of data their green line is better than their red line but you can actually like duplicate the effect of the paper which presumably was really hard for them to think of just by like adding more Teta right it's like why did they stop here and at like 200 thousand records right it'd probably be cheaper firm they'd go get like probably cheaper for their boss to go get like another two hundred thousand records then hire them to like do their advanced thing that they've got this lower so I mean obviously the Ruhr do you want to do both but I think it's also kind of telling you looking like where there's big advances in machine learning it's like always where there's like massive data sets right I mean like I think like people keep told me oh man like Facebook's like facial recognition algorithm is like so amazing it's like well of course like they have the most selfies they have so many selfies like they have like a billion labeled records of like faces like it makes sense of their the best facial recognition algorithm right like yeah they have the best people that do facial recognition and Facebook but I think they came after they got all that data and I would say if they like went to some other company that didn't have all that data I would still bet on Facebook having the best you know the best facial recognition algorithm and there's actually one more effect so I think a lot of people kind of agree with this these days I feel like two years ago is sort of showing this graph and as controversial or surprising but there's still kind of one effect that I think people do not study well enough and actually can't find any academic papers that really talk about this but there should be which is cleaner data right so if you if you have human collected human an annotated data it basically always will have errors in it and so you know one thing that I just simulated with my dataset was I added errors back into it right so I I look around built the model on percent accurate training data versus 95% accurate rainy versus like perfect training data and actually that had the same effect as collecting double and double the amounts of of data but it's pretty interesting because it's often easier to clean up your data set then collect twice as much new data so I actually think that that almost every company that we go into that I talk to like the best thing I could do for them is like just remind them of this right so like I think like anyone that's working with data that might be dirty it's definitely dirty and it probably probably the best thing you could do to make your algorithms work better is go through and root out those errors because those errors become the most important pieces of your data right because those are the things that surprise your algorithm so it ends up learning the most on them right so you know one thing you do is actually look at the training data that's affecting your algorithm most those are almost always mislabeled data right so I mean even like just look at the top 10 things that are like you know the kind of most important data records that you have often there's there's always something weird about them if your data set is sufficiently large and you should definitely look at them because the effect they can have on on algorithm is huge so you know I think that this is like this is sometimes surprising to machine learning people always surprising to machine learning people's bosses but I think that if you look at like kind of corporate strategies you can see they kind of agree with it right because like you see basically companies buy data and they open source their algorithms when I think that shows you what companies care about right they give away their algorithms and they they buy data right so you see IBM bought weather and Watson IBM bought merge healthcare for their data Facebook open source their learning tools Google open source or stuff Microsoft open sources their stuff you don't see any companies that give away their data and actually think that's a big problem we were just talking about this it's something that I I care about a lot I think it's like a you know data is such an advantage these days with machine learning that's become a huge issue for companies that can't get their hands on data right or or startups right that that you know need to compete they need to get this data so I mean one thing that I actually feel really proud of I mean there's lots of reasons to not like our government but I think like one thing that I really appreciate our governments doing is leading a movement for for open data right so they've open source I mean all these states actually have open data actually San Francisco's fantastic open data that's led to a lot a lot of great apps so it seems like there's a real push to open source except so you see like you know like learn sprout tracks educational outcomes would not be possible that open data street cred actually has a whole system to keep police officers safe like kind of predicting where there's going to be problems you see public art finder just helps citizens find like interesting public art none of these apps and these are all good useful apps they would not exist without the government opening up data right so I mean you hate to see the government kind of lead the private sector in terms of something useful right I mean Silicon Valley that seems like surprising but clearly that's happening right and you see Code for America in front of this so what are things that we wanted to do a crop flower that I always want to tell people about is we started to open up data and so you notice one of the data this is like very beginning of the company but I think this is a fun one we basically put like swaths of color and said what would you call this color and it led to this awesome cloud and there's only like one percent of the data that we collected for 50 bucks we've got so many labels because kind of a fun task and we took that that data set we put it online and I'll say by opening up the data it had the same good effects that open source has right because almost immediately some guy from sa P went through and he wants use the data so he basically cleans it up you fix all the misspellings remove the typos right and actually there's kind of a bigger effect here which is that depending on your application you might actually want the misspellings in the typos right so now we actually have two useful data sets right there's the original color data set and there's a cleaned up color data set both are useful depending on what you're trying to do and then people started to make really interesting visualizations off of off of this stuff right so here's a tag cloud where the size of the the colors is how many will set it right so you can see like green is like you know like just really top of mind for people right whereas like yeah there's like turquoise a smaller cerulean is like tiny right so you kind of see like what color is name English speakers you know think of and then and then data visualization tools start using it as an example because it's kind of a fun one and there's not many good open data sets that visualization tools can actually use to show off their stuff and then a kid made an app for colorblind people where you put a square around something that tells you not just like the red green and blue values of these colors but like what non colorblind people would call those colors right so saying you know this is like sky blue and this is navy blue right so kind of a rich experience for someone that they might be colorblind and so this worked worked super Welland and and we tried a couple times we did this thing with the Apple watch right where we were we were labeling tweets about the Apple watch and instead of kind of publishing our own you know results you thought well what if we just gave the raw data to data journalists so we let all journalists just become data journalists if they want to write because the datasets just kind of like handed to them right it's just like hey draw your own conclusions like whatever whatever you want from this and it was actually super cool right so like you know the you know what person does the most obvious thing and say okay what's the sentiment of this data set you know fifty-five percent of people said positive so that was like the lead for for this article but then you know someone from Texas just did one that was Texas related right awesome just like Texas doe like the Apple watch if you believe Twitter right which is totally accurate and interesting and then USA Today did one based on gender right so they said you know women love the Apple watch more than men and like the reason is because they think the applications are more useful than than men thing and so you know we were doing that and we were basically like collecting the same data over and over so this is also like sears out of date but like we did two years ago we've done 11,000 jobs that had sentiment in the title of the job right so we do tons and tons of sentiment collection because every job is like you know thousands are like tens of thousands of records that get labeled right so we're just like we keep labeling sentiment over and over and so we thought what if we just add an option where people could say that they'll make the data available to other people right so there's opening it up and you know we did a big launch around it and we actually made we may made a setting where you can set it and actually it got a fair amount in 2014 you know got a fair amount of results and you can use these right so we collected like 12 million hand label high quality data rows that you can use to build a sentiment classifier or lots of things and they're all available in our data for every one library I think it's specially like talking to meetups and kind of smaller companies we always really like I always want to tell you guys like you know I remember when we were just starting out and we you know we had no funding and no customers and no money right and like you know we love these like these public repositories so you know we would love for you guys to use it the license is super generous but then we did one more thing which you should also take advantage of which is you've made a data for everyone plan where if you're willing to make all your data publicly available which you know our bigger corporate customers are certainly not willing to do right so they have to pay license fees but if you are a start-up say or researcher and you are willing to make your data publicly available then we'll let you use our software for free and so that actually had a hockey stick effect on the popularity of opening up the data right so kind of giving people reason to open up the data cause us to get a lot more records this is actually continued I think we're I think we're over a billion data rows of open data that you can use and we have an API that you should totally hit because there's just like all kinds of interesting stuff in there and there's like URL categorization tasks so this is someone paid they paid over $7,000 to the crowd to label looks like almost a hundred thousand websites with the category of the websites in their taxonomy and you know they they made all that data available for it anyone that wants to build a website classifier you know there's all kinds of record data from PDFs which is probably boring for a lot of applications but you know if you need to build a OCR system or a text extraction system this would be a fantastic way to test it you know there's name and title extraction which is why these things that seems really simple until you in you've actually worked on it alright so this is like you know saying if it's dr. Lucas be walled that my first name isn't dr. and summarization even weird stuff like isn't image funny I mean this is uh this is probably this may be the only person that ever wants to use this but I've had I've seen some machine learning classes have tried to build this classifiers like a particularly tough challenge for for deep learning you're more serious about classifying medical images attributes of people you know anyway all these data sets are available and and we'd absolutely love for for you guys to use it so thank you sure oh sure so the question is that is is basically everybody has kind of a different tax on your different way they want to classify data and kind of what are we doing about that and I would say we haven't looked deeply into it I mean I think you're absolutely right that everyone has a different taxonomy and it's incredibly good for crop flowers business that everyone has a different taxonomy I I think um I think that I guess I guess it's to be honest I don't really have smart things to say about it I mean it's it's like I think that I would say that here's what I really say is like when when you can use your own taxonomy it's incredibly freeing and helpful I think there's a lot of work on trying to like take data that's classified one way and like you use that data for a different type of classification right so there's like you know I mean transfer learning in a way is like a type of you know thing that kind of leans that way I think that that's only gonna become more and more important because it kind of what we've seen from our customers is that they really insist on their taxonomy so like you know we work with like you know Home Depot and Lowe's right and they have such similar inventory but they're never gonna collaborate on a taxonomy of hardware you know cuz it's just like so fundamental to like you know what they're trying to do or like you know even like setting aside taxonomy it's like you know what's the definition of an address of a business right so so many of our customers are like I wanted the address of the business right and I don't think they even realize like how underspecified that is right so if like you know decide like what's it what is the address of McDonald's right it depends on what you're trying to do right if you're trying to make a map you might want the McDonald's that's like next door to here right if you're trying to like sell McDonald's some software like you might want there like international headquarters if you know you want you know you might want like a mailing address I mean it's just like they're there's so many different ways that that that that can mean I think it's a I just think it's incredibly hard to figure out how to how to combine them so I guess I don't have any genius way to do it but if you do you should right right ya know it's a good point I mean I think the problem is a deal taxonomy depends so much on what you're doing if we don't know what somebody's kind of doing downstream from us it's hard for us to say like you know we can say like what's the taxonomy that people can understand and classify but that might not be what somebody actually wants but but yeah I mean I think this stuff is really interesting yeah why sure sure so the question is around the mirror who are the kind of people doing the the work on the back end and I would say you know we started off as a very pure you know kind of open platform where anyone could log in and start earning money and and kind of like an open marketplace where people would just post prices and people would log in and we got all types of different people doing that so so it was a really broad demographic I'm actually published a bunch of studies in the demographics but you know it turns out it's like you know Half Men half women that kind of follows that you know if you look at by state it follows a population of states it's about half the folks in the US have a college degree but then you know we had a lot of customers that said hey you know we care a lot about the demographic distributions of the killer' about the demographics of the people that are doing these tasks in the back end so what we did was actually partnered with a whole bunch of outsourcing companies and we basically say if you want to send this to a particular outsourcing company or even a particular group of people like your own employees you can do that right so you can if you want to send it to you know if you want to send to a particular company in India where people are just like doing this labeling all day long you can totally do that if you want to send it to like veterans you can you can do that and if you want to if you want just like a link that you can pass around you know among your your team and just like you know say hey can everyone like label 100 of these you can also do that so so I think it's like absolutely true that the the demographics of people matters a ton but it's maybe sort of similar to the taxonomy question that it matters a ton but there's not like a right solution for like any particular person it kind of depends on what and what you're trying to do like you know if you're trying to get you know if you're trying to label if you're trying to if you're like a fashion company that wants to like label all the dresses in your inventory you probably want like a really different kind of person then you know if your home depot trying to label like all the you know all the hardware inventory and and that just like that would almost look like the exact same task right so because that's just sort of like categorizing inventory so we kind of rely on our customers and our to use our tool to get the right demographics to get the results they want yeah I'm curious like here at home so I guess the question is you know I'm talking a lot about kind of like classification and maybe not like kind of like prediction of the future you know okay I'm like I think I think I think people often talk about classification as you know because it's kind of this the simplest application of machine learning I think I think in practice things get a lot more complicated you know I think I mean the way I would look at at prediction is sort of you know kind of taking you know taking all the data you have up to like a certain time and then you know trying to trying to classify or trying to do a regression on you know the next set of data that I have like that's how I would kind of frame the the friction problem but you know actually there's probably like a lot a lot of folks here that are you know kind of more qualified than me to talk about you know prediction in general so I you know I'm not sure that I think I think I would say I think a lot of customers a lot of our customers are doing things that you'd probably call prediction that they might call classification and and vice versa because it's there's a lot of overlap but in terms of like time series analysis I mean I guess some of our customers do it but it's not my area of expertise so probably I don't have too much smart to say but yeah right yeah I mean I think like you should protect Lee ask all the people that raised their hand saying they work on machine learning here cuz it was like almost everyone I think like one thing that's frustrating about AI is the language is so unnecessarily confusing and you know I think a lot of things that are simple get kind of highfalutin terms attached to them and so I think that's terrible so I think I quinces yeah I would just ask ask everyone here yeah maybe maybe you guys can help out this this man over here anyone just an a prediction if I don't feel like I should let let let you guys see the next talk and then go home Thanks