Devreal

Text By the Bay 2015: John Akred & Rob Monroe, World’s Largest Database of Car Features from PDFs

Text By the Bay 2015: John Akred & Rob Monroe, World’s Largest Database of Car Features from PDFs

Recording: Text By the Bay 2015: John Akred & Rob Monroe, World’s Largest Database of Car Features from PDFs

good afternoon everybody um hopefully you can can everybody hear me okay excellent all right we won't worry so much about that one we're going to be passing this what will seem like a cosmetic microphone to you but it's for the benefit of the camera um so i'm john acrid this is rob monroe we'll tell you a little bit more about ourselves in just a second but we're going to talk about some work we did together with both of our companies uh for a client of ours edmunds.com which you may be familiar with as a car research and shopping site as you can verify that is me so i'm the cto of a consulting company based in mountain view we're about 50 people and we work with um we do fun things with data for our clients typically solving hard problems for which there's no out-of-the-box solution to do it all often working with fantastic startups who bring new technologies to market to make those problems solvable in the first place and this is rob on the microphones so hi everyone i'm robert monroe i'm the ceo of idibom where a technology company based here in san francisco and we specialize in text analytics into particular dimensions uh one is that we work in any language um so we worked in more than 50 languages to date and we try to get to state-of-the-art accuracy across a large number of languages the other side and this will become particularly relevant here in this talk is that we specialize in integrating natural language processing with human annotation in order to get to high levels of accuracy as quickly as possible so we have um two members of eddie barn here in the audience as well um just um so a tripty in genre if you can your hands up and say hi to everyone um we're really interested in hearing and about any kind of problems that people are working on in the tax analytics space so we have clients which range from the most visited car website in the world to the united nations so we help unicef process text messages sent by citizens and populations worldwide um so always keen to hear more but today we're we're talking about the auto industry how's it going so i'm going to start by talking a little bit about why we did this project in the first place um it's going to sound a little bit like a data strategy interlude but it also highlights i think the opportunity with linguistic assets that can be applied across a range of enterprise use cases not just the one we're going to talk about today which is which is sort of why we ended up choosing this project but we work with customers to define their road maps of technology investments and when we sat down with edmunds their problem wasn't just the one we're going to talk about today they they work in a business that is swimming in a lot of unstructured data so they are the world's most popular car site they have folks that provide commentary about those cars and what they like about them and what they don't and things like that they do full-on editorial articles about cars and reviews and things like that and and then there's the the descriptions um of the cars themselves is just a few examples of where there's unstructured data in their enterprise so as we sat down with them to to think about um what projects we might do to have the most impact on their business quickly it became obvious that improving their ability to work with text and unstructured data was critical and and the way we sort of arrive at that conclusion is we we tend to look across various workloads in an organization um so if if you can think of maybe let's call use case one the the browsing for information about cars themselves research if you will and let's call use case two then looking for individual car inventories going from i'm looking at the bmw four series two hey there's this red one with this engine that looks like it's it's the car meant for me on this dealer's lot uh so let's call that use case number two uh and maybe use case number three is the customer comments about cars uh that they accept and allow people to have conversations on their forum right um workload b you'll notice is is common to all of those and so we would we would think about the kind of use case we're going to describe as one that that we're going to talk about a specific instance but we can take the assets that we built in the process of doing this the ontology the ontological resources the classifiers the models that have had the benefit of lots of review and training and point them at different use cases in the organization so one of the powerful things about text analytics capabilities when architected well is there's typically a bunch of things in an organization you can point them at and oftentimes that argument is what gets those things rising in the pile of possible things that a group and an organization might tackle in terms of technology innovation uh so we sat down we thought about what their priorities were we looked at the dimensions of of what they were trying to do in their business and their assumptions and in this case that means that they wanted to go to a world of understanding individual car inventory at a very deep level that right now they had one way of doing that which was uh by manually processing the the pdfs you'll see shortly that hold the information about cars and what their features and options are and and their assumption was that there was well up until we we engaged in this conversation was there were no technical assets that could automate or at least semi-automate and speed that process for them and so you know this is a situation where you've got a combination of a startup who's bringing a product to market that allows you to process text and do the various fun things like entity resolution concept discovery etc with it uh in in a way that opens up a problem space to an enterprise like edmonds over which they have several uses and so as we go through this we we we never think of these things as static technology is always changing right so um simultaneously the the uh the conclusions we we came out of the first set of conversations with said hey let's go do this stuff um as their business matures as as they come to do other things the prioritization may change and actually one of the interesting things is we've we've sort of revisited this original conversation with them three or four times since we started the journey we're about to describe uh and and it's it's interesting that the text analytics has stayed right at the top of their uh their priorities and and uh along with some some other of the conversation so um i think uh at a conference like this i'm probably preaching to the choir but but the the the fact that you've got an organization uh who has so many uses for this kind of technology and capability i think speaks to what's happening in the market around these skills and opportunities as well so let's talk about edmunds the the world's um most visited site about cars they actually almost go back this far they they had a gopher site does anybody remember gopher yes awesome precursor to mosaic i think from the university of minnesota because they're the gophers um but at any rate uh so so very technologically savvy organization that likes to work with startups uh and explore new technologies not all organizations in the world are like this but but we like them because they're they're very keen to experiment uh with new capabilities to to drive their own competitive advantage uh and and this is a story of that so to set up the the problem that that we had um you know you see here that they've got a website and um typically people start out by looking at what i will call classes of cars and i mean that in the object oriented sense not like the mercedes c-class sense which occurs me i just had a name collision there um but but uh you know you're looking for i want a two-door i want a coupe you know with a sporty engine or something right and and mercedes bmw all these places have an answer to that but um but if you're looking for one with four-wheel drive you know mercedes calls that 4matic audi calls it quattro uh bmw calls it x drive everybody's got their own special term right so um you can already see sort of some of the linguistic and and uh text analysis challenges that the car industry presents um rob and i actually met at kdd uh a couple two and a half years ago i guess at this point um in chicago and and over beers i was i was uh commiserating about this this thorny problem and rob happened to unbeknownst to me at the time have a startup who was focused on car dealer data or car data at the time so uh like like all great stories this idea was born in a bar so once you go through this this research experience of looking at different models and the kind of features and options that are available to you you typically want to then move to okay i think i found my you know the twinkle in my eye around this one car it's funny giving this talk in a city full of people who don't actually drive cars but um but but down in l.a they sure seem fun there is a world out there where people drive um i occasionally have to drive to the train um but uh the you know they they they they quickly want to find that the inventory of the kind of car they've discovered and start to understand it and it turns out that that descriptions of inventory are very very noisy and the dealer can do a lot to a car after it arrives on the lot so they have their dealer added options and and one thing that they want to get right is get a very very solid description of these actual cars and how they're configured so that you don't go oh awesome the one i've been waiting for all my life is over there you go to the dealer and it's like nope doesn't have that or something right so there's a high cost of misclassifications i guess what i'm saying um so they wanted to go from a site that was primarily interested in the class of vehicle right a 2015 c-class to that serial number one two three four 2015 c-class that's black with that engine air-conditioned dual front zone air conditioning etc and the way they get the information that describes that looks like this can everyone read that just be thankful you never have to so so before rob and i came along um edmonds had a team of of about i wanna say about 40 people uh data editors whose job it was to take and this is an excerpt of what's probably a few hundred page document of it's a you know it's a german car company so they in exquisite detail describe each and every option and feature and the rules about which ones can co-occur and all that kind of fun stuff right um great if you want to sit down and night when you have insomnia uh and and browse a little bit about things you never knew about the the four series sorry the five series uh but if you want to populate a structured database of car features and options to support faceted search across models kind of a bad format i describe pdfs as a place where good data goes to die don't see i'm paraphrasing john rouser who once described the enterprise data warehouse thusly and i also agree with that um but the the um every it seems like every two years or so i get a project where the data is hiding in some pdfs and i optimistically am like okay maybe this time someone's got it and and uh every time and this is this is not an insult to the people who are trying to solve this problem it's uh uh admiration of the complexity of what adobe has created with the pdf spec i suppose but it always presents a gigantic problem and and i know some folks around here are engaged in making that problem better and we wish you all the luck in the world there's nothing we would like better than to not have to go through what i'm about to describe but that said um the problem with this world was it would take a human about three weeks to basically hand enter this information into a database that backed the the faceted search and inventory on their site and the the the release of new models is not evenly distributed throughout the year all the car manufacturers tend to do it well earlier and earlier but sometime in late summer you start hearing about the next year's models and so you know on average that means about seven percent of the world of the of the available cars in any region don't make it up onto their site as inventory and if you've got a new model that that represents an entire class of cars don't get up on their site until some human goes through this process now if you think about a model where you're advertising um supported so you've got a site that describes a car bmw releases the four series which a year or two ago was a new model and people are curious about that when they read the announcement and they find out more and the interest in that drops off over time you would imagine right so if you can actually shift forward in time how quickly they can get to that structured understanding and have that vehicle find it when somebody's searching for a two-door coupe you're going to make a lot more money on advertising revenue on the one hand and then like i described you also want to get all that inventory discoverable because if they pay some small small amount for an ad a physical lead of an interested car buyer coming from a site who's done this search found the car gone to a list and clicked and said i want a quote on that one is much much much more valuable it's sense versus 50 bucks or something like that so so having six and a half percent of all the available inventory in a market undiscoverable is actually leaving a lot a lot of money on the table and this is also a very siloed operation so so this is what we set out to fix um we are working with an existing human process so one of the weight reasons we matched so well with an 80 solution as rob described as they specialize on a human feedback loop around what the math is doing and we had this is this is as a consultant this is incredibly refreshing we worked with a company who wanted to enable a team of humans with technology not just so they could have fewer of them and fire a bunch usually when you're automating or semi-automating someone's job the goal is actually to automate that workforce out of existence and i'm happy to say that in this case it was actually hey there's all these other things that those people could be looking for and there's a lot of value we could extract in our organization if we can empower them and take care of the 80th percent of the stuff that's that's easily automatable so we worked with the the folks on the one hand to sort of adjust the way they looked at products go from a strict engineered system to a probabilistic one if you will right so that um rather than having to have the full structured model of a car on the website you can you can do some fuzzy matching in those search results and actually um change the nature of that time to market problem by making it less uh having to have less in place to get it to market and then the next step also make it much faster to do that part of that that you still have to do um one of our big takeaways is that that um you have to build this into the process right for it to work well we we fortunately anticipated that obviously uh the whole goal was to empower this team so we knew you're gonna have to work very closely with them uh to enable them to do their thing but um it's it's often i think an overlooked point when folks are building systems with with fancy machine learning and nlp capabilities that that last mile of how you surface that capability to the people who are ultimately going to be using it is really really important so this is what we built a nice simple picture of boxes and arrows and i will describe the flow briefly and then we'll get a bit more deep into the the actual technology and how we did it but but to set the stage for that deeper discussion we pull in those pdfs and we use scraper scraper wiki in this case to extract the basic data within that pdf we had the benefit of having humans involved in this so the the extraction did not have to be perfect and as you may recall from a minute ago it's as far as i'm aware nobody's got a perfect pdf extractor yet um so so the editors take that initial um extraction coach it somewhat to make sure that it's it's it's reading the documents correctly before importing that extracted information in the form of a csv uh into the data environment we use to do do the actual work once in there we've basically got um a database that has the the source vehicle information and what happens and we'll show the actual user innovation components of this but there's a data editor who is basically sending descriptive chunks of texts things like dual zone front air conditioning to the idibon service which is giving back a prediction on what actual ontological element that refers to a semantic element so there's lots of ways to describe air conditioning this is that class of dual zone front air conditioning the canonical thing not the description i just used to tell you what it is and and then over time they set certain thresholds to say well if the classification is confident enough i will accept that and uh if not i will label it and correct it and so it goes it has a feedback loop so as we come across things uh that the classifier gets wrong we can coach it and then future subsequent sorry subsequent predictions will take advantage of that labeling interestingly as we were working on this project an option showed up well an item showed up that was vacuum cleaner we were like oh is that a typo is there a vacuum cleaner in a car turns out honda put a vacuum cleaner in a minivan if you have a minivan and put pets or children in it regularly you probably understand this i don't have a minivan i do have dogs waiting for them to put vacuum cleaners and smaller models uh but but turns out that's that's a classic outlier right and so the system is great because when something like that shows up the the the editor is able to process it thusly so this is sort of the existing tool that the edmonds editors used it's very cheap and cheerful sort of windows you know style tree navigation there you know you can see that there's 2015 there's some codes there that don't mean a whole lot but ultimately you can see we're in a series of bmw buckets here and then these are the individual trim packages so for them a class of vehicle is at the trim level right so the the the um 528i with x drive yada yada has a has its own is distinct from say the 535 without four wheel drive and what they do is they they go and find the model or there's if this is a new model they would create a new bucket and then they would start importing from csv which is what fires up the whole importation process i described and they get something like this which if you think back to that you know seems a pretty reasonable rec representation of a table of features and options that that was hard to read um it actually is pulling through symbols um in the boxes and most of these documents have a bizarre language of symbology of what they put in the boxes which is is you know typically one of those is yes one of those is maybe and one of those is no but this is sort of the the first extraction stage so as long as it's got the columns right um you know we sort of annotate some stuff up top tell it where the real data is um then we can get start surfacing this information to the capability that rob's going to describe in a way that allows us to use the prediction service shoving a 300 page pdf at rob service might have some humorous results but it's not designed to to to parse apart a document it's designed to help you do the kinds of things that we're about to describe but so the chunking up of the problem is always an interesting part of the whole text analysis process so you can see they start um saying where things are and and defining the columns to to the extractor and ultimately wind up with a in a world where those entities which are the columns on the left have been successfully extracted and what we're trying to do is guess the three-level hierarchy in the middle which i'm guessing is hard to read so i will describe it a little bit there are a thing the the levels of the hierarchy are attribute group so that's things like uh the draw stuff about the driver's seat or um you know stuff about airbags uh and within that we have an attribute name so if it's the class is airbags and you've got you know these these cars have lots of airbags these days so one of them might be the head airbag and then the attribute value might be front or rear so you're getting to a point where you're saying there's an airbag it's a head airbag and it's in the front seat and that is the level of specificity that these features and options are described or represented i should say in edmond's internal database to support the kind of faceted search and discovery and then over on the right you can see that in this interface they basically ask for a prediction and then when they've got it they get a confidence interval here that's sort of the combined confidence of that prediction rob's going to talk a lot more about how we do this in depth but we take advantage of the con of the hierarchy that we're trying to predict by sort of stepping through it so we get a prediction on attribute group when we can narrow it to air conditioning obviously we can uh we're essentially narrowing the problem space of the text we're trying to classify now to within the air conditioning category um and and in this case maybe it's air filtration and then the value is yeah there's a there's active charcoal filtering in the uh the the air condition system which if you have a puppy like i do is a good thing so that's sort of how we framed the the problem to the actual nlp capability any questions before uh this is the phrase in the source document that for instance bmw's provides to describe the car so bmw says we have a feature called micro filter ventilation system with replaceable active charcoal filters now um although it's in the the interesting thing is it's an internal sourcing document so if anything it's meant to bedazzle the buyer at the dealer who's picking how many cars of what varieties they're gonna ask for so it's probably reused marketing materials if they actually used more sober descriptions of these things this problem might be easier but that's sort of what's in that giant pdf in the in the row and then it ticks off which option packages include it and things like that which is a part of a downstream part of the problem that we're trying to solve so we we feed that chunk of text once we parsed it out of the pdf into rob's classifier which he is now going to describe before we took on this project so there are five kinds of head rests like who can name two right um that counts so um so you can see the uh the level of specificity there so i mean it's not just air conditioning it's not just filtration it's a particular type of filtration so in total i think that we have the number coming up so i think it's about four thousand um you could think of them as entities in total um or a classification task so classifying the short chunks of text into one or four thousand categories um and as many of you know four thousand categories is a lot for an nlp task uh typically you'd want a large amount of training data um much more than we're actually gonna have so i'll talk about the data sparsity issue and how it fits in the ontology and how it actually turned out uh not to be a big problem for us i think it's it's interesting and fairly unique um but first of all i want to talk about the uh our fundamental approach so like john said we specialize in that feedback loop between humans and machine learning we certainly weren't trying to um replace admins analysts we're trying to make them more efficient so we're trying to reduce the the two to three weeks they're you to get a new car onto their website uh to to um have that processed in as short a time as possible uh to remove the the redundant work from the analysts so it really is only the the novel or the inherently ambiguous items that they have to review so uh like i said it's not just empowering the analyst it's actually a better approach to more accurate natural language processing as well so if you saw the keynote yesterday mark lieberman was very happy that almost every single paper at the last acl was using an existing open data set and comparing themselves to the previous results i actually found that a little bit depressing because this very academic model of keeping the data constant and modifying the algorithm does not necessarily give you the most accurate data so here's a really good example so imagine that you're trying to disambiguate the word forward in this case this is data taken from uh social media across a few languages um so the ford mustang would be the correct user for this is the organization or the car use um but certainly you know imagine you're tracking the ford brand on social media um you don't want rob ford's latest gaff in toronto to throw off your sentiment analysis um and you don't want harrison ford's new movie or lego character um to um to again like throw off your analysis uh so this disambiguation is um uh either something we do directly or is built into a lot of our tasks lots of brands lots of things related to automobile in particular are ambiguous terms or terms that we'll also use in everyday life so if we apply a linear model with just some you know fairly standard features like words and grams shapes of features um i think we use dictionaries in this case as well um then we get uh 0.45 f value so you know almost 50 accuracy in distinguishing whether it's for the automobile um from whether it is a person use afford this is just the binary organization versus um um person usage uh i don't think there are any or not very many other sensors afford in this data so if we switch to deep learning algorithms we get about a two percent gain in accuracy so a two percent gain in named any recognition will you know get you a maybe a couple of publications um at acl certainly it's enough in itself that'll get you a phd um it's not enough that if you go to someone in industry and say hey we're two percent better than the competition that they're going to fall over themselves unfortunately um um maybe on wall street right yeah yeah i mean yeah if you're in a hedge fund and you're doing high frequency trades yeah there are use cases so um in domain training so simply take in data um which is from the same domain so data specifically around the auto industry and applying this to this particular task uh then you get this huge jump um so you're at about um uh 0.61.62 uh in f value um but uh using idibon if we give um the right items uh to our annotators i'm not going to go into too much detail about what the ride items mean but you can imagine it's a combination of items which are ambiguous items which promote greater coverage etc just 10 minutes of nls feedback and we're automating this task at almost 95 accuracy so this is just the automated analysis obviously with the humans on top we're closer to 100. so these are the kinds of gains you get by optimizing humans in the loop in a really really smart way so 10 minutes of nls feedback you know versus you know a phd um uh so um uh yeah so this uh speaks very much to our value proposition um as well as to this particular use case so admins ontology um has four thousand different car features uh the their actual ontology is proprietary data so we've shown you just a small part of it here and so for this uh the sentence that you saw microfilter ventilation systems with replaceable active charcoal filters just rolls off the tongue right so this is part of the ontology where it could fit in so at the top level i condition in the next level we've got things like filtration climate control memory front and rear and then with infiltration there's a number of categories in them including the one that you saw here um so uh we have about four thousand features in this ontology in total um so i think it's something like more than three thousand five hundred leaf nodes and um there's only uh twenty thousand world called label items so existing data points for which the analysts at admins have said yes this belongs in in this particular category uh so this means it's really sparse data right uh so on average you've only got um five labels per um five indicated items per label um but we need to provide a system which will requires the input of these expertise is also accurate with very few data points so again we can't take the academic position of wait until we've got 10 000 items and then cross-validating on that we have to be smart about the ways we get to accuracy quickly so at the the top levels yeah so the level of air conditioning we probably have several hundred items at the very highest levels whether it's actually a car feature or not that's when you've got um thousands uh so the features we use um words sequences of words engrams um some existing technology taxonomies and word shapes so if you're not feeling that particular expression you can think of word shape as a form of industry specific regular expressions so you can imagine that a feature might say any number of um uh numbers followed by l so you're like a five liter engine a ten liter engine etc um and so that allows you to generalize over some of the terminology that that might be expressed in in very different ways um negative sampling was interesting as well so we're mainly getting through positive items here are the features um but we're trained in here on a regular classification survives classification tasks so we need those negative items as well oh i signed the things and i wouldn't walk out of camera am i walking out of camera right now what about now sorry sorry lexi all right um so um and so in this case we would um negatively sample items within the same parent so if we're looking for um air filtration we only have the positive examples there we know that um in this particular case it's not mutually exclusive because we haven't seen the same item uh co-labeled with two labels which they can have then we can pull the negative labels from front and rear air conditioning knowing that by the time the classification has come down from air conditioning we're only doing the um uh the discrimination task within that that subcategory uh so there's the inherent negative labels for our training items um and then in the complete absence of all data um and this is actually really important from user point of view someone types in this new label and they say hey i've got this new label interior active charcoal filter um that from your user perspective should be enough that you can start making reasonable guesses about that right so if you have a label in there called interior active charcoal filter and then you've got this micro filter ventilation sort of placebo alcohol checked chuckle filters up top there is actually information there so at that point before we even have any labeled items which have come back from the admins analysts uh we treat the the label itself as the um as the data point from which we extract the text um and this was a really minor thing that seems obvious when you talk about it and haven't seen anyone else do this before maybe you've had to do this in in tasks in the past but it was something that kind of overnight really changed the user experience not because i was particularly confident with just one data point no matter how much you oversample more because it provides some reasonable feedback even at low accuracy which felt responsive to the end users so um something i think is really interesting that that falls out of this is the accuracy at different depths of the ontology um so cross-validation accuracy broadly you expect accuracy to go up with more raw data items that's one of our one of our core value props um but when we look at the accuracy the average accuracy at different depths of the ontology it actually goes up so if we're at the first depth uh we're right now getting you know around about 0.8 in in f value in terms of accuracy of classifying you know say air conditioning versus head rests um the um the precision is quite high which in this particular use case is is useful for us um if we go to depth two so discriminating air filtration from climate control memory or front and rear zones um then the actually is actually a little bit higher it's 0.81 and the precision is again pushing about 0.9 and then if we go to the the third depth this is where we often only have a handful of labeled items accuracy is actually really really high um it's 0.929 so this is like approaching the level of inter-annotator agreement for a lot of these classes so is there a dependency between currently classifying depth one and um there is in ultimate predictions yeah um so that's something um that admins take care of their sites actually in the slides here but yeah so they'll start at the top level um building any one system is that can cascade by confidence admins actually do something a little bit more sophisticated in that they use a combination of confidence at the top level um and the cross-validation accuracy at the the top level um they use both of those programmatically to optimize what then uh gets populated further down that ontology for prediction that's right that's right yeah yeah so the average branching factor is something around five uh typically yeah um yeah and you can see the precision is quite high here as well um not shown here as well though um at this depth um a lot of items will come out at lower confidence um so this is something interesting that we're not trying to tackle right now and the system is working but it's something to think about like well if we know that we're actually higher at this level of of accuracy but our confidence isn't up there simply because we have fewer data items that the lower confidence is mostly the result of smoothing um you know other smart ways uh that we can sensibly adjust that confidence when we have those uh smaller number of training items uh if anyone from the more technical side has some ideas about that i'd love to hear about it i came up in theoretical nlp i got my phd in the area of stanford before i started running a business so i'm really happy to geek out on the on the ml side of things um and uh the reason that we're getting this that this accuracy even though it's low confidence is that it's controlled vocabulary right so at the height the very high levels you've got things like like forward right um it's inherently ambiguous it's just a tough task for for nlp uh same with their conditioning um so something like aircon you know can be abbreviated by things just like ac right so ac is inherently ambiguous acronym um that can map to a number of other things um although i can't think of anything right now anyone ac alternating current oh right ac dc one of my favorite bands all right thank you um yeah um whereas um i don't even know if there's an acronym or another way of expressing interior active charcoal filters but um if if iacf is something else it's probably not that widely used um so we uh we get the advantage of at the higher levels like of saying well even it's it's even related to cars we're at the lower level we've kind of gotten rid of that that relevance question and so it really is about the specific parts of the ontology anyway so i thought this was a really interesting insight it's something that we don't typically see in nlp um so we hand this over to john so um as we mentioned it was taken two to three weeks from the world's most visited car site to get information about new cars onto their site when everyone was searching for them and john's going to tell us how long it's taken them now so their backlog is massively reduced and this is all done via apis so um so that they can extend this to other areas um like i said the a lot of this is setting it up well to support the the existing workflow and processes uh to realize that benefit um and and that requires some coaching of the folks and making that easy for them so the the point i'll leave with is that this asset to the to how we sort of kick this off from a strategy perspective allows you to use what we've developed by training a bunch of of classifiers and and nlp capabilities in one task and point that capability at things like comments and do things like entity resolution and concept discovery to understand the different way people refer to these things and ultimately enable sentiment analysis on comments so what what i think is a really nice story and it's on its own merits of of effectively leveraging a very powerful nlp capability to solve a real business problem um is is even more goodness when you think about the fact that most organizations have several use cases that they can put these things to so with that we'll thank you and as rob said we're happy to geek out on these things so the today's turnaround time what's the percentage of a human what's the percentage of the human in loop i think um not just because it's a famous rule we're in the neighborhood of 75 to 80 percent of the features are accurately predicted i don't know where they've set the threshold within their ui i should check with should have checked with them before we spoke but in other words when we were last talking you know somewhere up in the 85 or 90 confidence range is where they would started letting actually taking the human out of the loop and just saying okay we're going to go with that prediction um they they they are on the one hand continuously training the things so that their predictions on average over time are going to get more confident on the other and they are getting more confident with that tool and may actually set that threshold um slightly lower over time yeah and even the humans are in the loop like like john demonstrated that's a lot more efficient yeah rather than scrolling through the page 30 and typing something out it's having that immediately presented to them for review and they don't seem to miss reclassifying the same boring thing over and over and over again thanks everybody you