data.bythebay.io: Richard Downe - Hidden in plain sight: Using law to summarize the law
Recording: data.bythebay.io: Richard Downe - Hidden in plain sight: Using law to summarize the law
everybody thank you for coming to my talk I'll just get started so I don't know if you're not familiar with case text I do have a couple of slides to just introduce us I think they put the illegal research companies right stacked in a roach is good get the audience in a room and where everybody out so we are CEO went through Y Combinator he's an ex litigator as are a number of our other members of the company including our general counsel and the kind of the inspiration for it was the dissatisfaction with the tools that were available the fact they were so horribly expensive and not actually that pleasant to use so you're probably familiar with the market the goal here is to provide a better alternative we believe we can do what West and Lexus do better faster cheaper so I not sure how you know many in the audience are technical versus lawyers but the most important thing when you're arguing cases you can find relevant precedent and you can explain why your argument built on previous cases is going to prevail in court over the other guys and so you know one thing where technology becomes extremely critical as time goes on you have more precedent you have more law there's more you'd have to wade through and read through so without technology it would just be more time reading through books in a library which is on its face fun but if you've ever tried to read a judicial opinion it's actually not at any rate we have first thing I want to do is give you an overview on our actual architecture at case text so we have kind of evolved quite a bit when the site started out in the Y Combinator today's we were just a django site we built a batch processing system about two years ago we used a technology called hgtv forwarding that is not used anywhere else in the universe and it had its advantages it was simple it was easy to build every library in every language every language in the world has a library support it that was part of the appeal was that you know if you wanted to write something in C++ or go or whatever you could sit down and put a little piece of code in for it of course there's plenty of other companies trying to build pipeline technology and we're not an infrastructure company so in the past year I have moved everything over to spark in DMR we do not have any ops people in the company which means that I get to wear an ops hat when we need to build stuff so if I go with everything has been to minimize the amount of actual like maintenance and babysitting at all requires EMR is fantastic because I don't have to be a Hadoop administrator ever you know when I was learning spark the one thing I learned talking to everyone who used it was the people who really like to do all had big companies with dev ops teams where the people who actually like to do never actually had to worry about installing it or configuring it or setting it up and thankfully Amazon does the hard work for us one of the most important pieces at the beginning of this hopper is XML transformation it's also the least glamorous but we pull in vendor sourced XML files for the judicial opinions most statutes and regulations were able to get actually directly from the government there's a gentleman by the name of Aria her shoe it's who works with the library of congress to produce a very very high quality xml version of the u.s. code which we use on our site and the CFR the Code of Federal Regulations is not quite as nice they have printer typesetting coats in the XML and that's quite a bit more work to get but it any rate so we have a bunch of data munching that transforms all these XML files into a common format and that gives us one single format that all of our other code has to work with so the fun stuff the NLP analysis Apache weima is a library that was originally put together for the Watson jeopardy project it since has found adoption all around world thanks to being an Apache project and of course as things go and big companies there is now a faction in outside IBM that is trying to replace it with an alternative technology because that's what happens when you have as many people in your company is a small country I spent a number of years at IBM and actually the I was introduced to we working on analyzing medical literature at the element in lab in south san jose and then when I went to case text I thought you know this is a useful tool for the law as well any rate what's really cool about it is that you get this index that you can navigate for all the different entities you find so it doesn't actually give you any particular facility for identifying stuff but what it gives you is a great tool for actually manipulating the different things together so we use a variety of techniques for actually finding things like citations case titles inside of the text now a title is a difficult thing because it's there's a lot of ambiguity and that has the title edit or not it takes traditional any our problems and kind of turns them up a notch because you are trying to find a juxtaposition of entities and you're trying to find where their actual description as a joint entity ends versus being discussed in reference to some other entity and so the fact that you're looking for a concatenation of entities is tricky thankfully most case titles I think something like ninety-eight percent of all case titles have a video in middle a coordinating conjunction to use the part of speech raising from the Stanford tiger and so that actually makes it quite easy to find but the point is that if you are able to find sort of ancillary information you can improve your odds of actually getting a match dramatically so you want to find a docket number you want to find a date of the decision now you have a title a jacket a date and a jurisdiction you have a chance of correctly identifying the title of the case otherwise you know there are many john vedoe in fact there are i think something in the neighborhood of 500 cases that are i think some particular insurance company versus state there's there we have some very litigious insurance companies and they have sued the government many many times so this is kind of pipeline overload we have you know we may pipelines running inside of spark pipelines all manipulated by air flow which is more pipelines but that seems to be the way things go these days the architecture is you know simple in broad strokes data lives in s3 except when it is relational and it lives in Postgres the goal is to minimize the amount of data that actually has to live in a relational database but case does have relational aspects to it so it is nice to be able to do that postgres is a fantastic tool and what goes in there is really only a couple of gigabytes of metadata all the real stuff lives in s3 and the goal with the construction of this system has been to make sure that very few operations ever have to actually touch postgres it's kind of used as a kind of a baseline of storing that these XML files are all actually different copies of the same case and giving us some ability to say well if prefer the one that's published one that's from the better vendor the one that we didn't scrape from the court itself for a while when we first started out I actually wrote a little thing that took the PDF files from the Federal Circuit every morning just stripped out the text and used a bunch of regular expressions on the white space to get decent HTML out of it this is because we were trying to convince patent litigators that we were a irrelevant alternative to west and lexus this is you know in the very early days when we didn't even have state courts online so it was any niche that we could get easily in a timely manner one thing that's great about working with law is that it actually is very in terminus to most English text as you know is not something you can analyze with a you know context-free grammar but the law actually is this is why reading the opinions is not the most you know joyous thing in the world but it all is very very funny like the blue book is the style guide used by lawyers and it specifies in great detail what everything has to look like and the great thing about that is it's an opportunity to find lots of low-hanging fruit before you start delving into the mls I technologies and it was great about that is when you have found tons and tons of stuff with really high confidence using deterministic methods you have an enormous amount of kind of scaffolding to build an ml approach on you know where the citations are you know what the great citation graph looks like you know where your jurisdictions are so in terms of figuring out how the cases relate to each other you're already miles ahead than if you were just trying to use it based upon word to Veck or word soup or any other type of more general technique so these are anyway this is so these summaries the parenthetical summaries are when one judge quotes another or sites of the case and actually see if I can find an example of one right now so here is it brown v board of education on our website and screws in you all right if I can do this without again okay so you see at the very top we see a few of these summaries that we've harvested from other cases and these are really they're fantastic they really do summarize what we understand Brown Viva were to be about and if we click on one of these links we can go and see you where it is in context so you see here's one of them it's not the brown v board one but I did it yeah so here we can see no this is a completely different one well this is yeah but any rate you see an example of them in context and so you see the citation then you see the parenthetical and the judge does very succinctly summarized what the case that is being cited is about so in terms of presenting context to our users and presenting context of lawyers this actually is fantastic for at the very top of the case being able to say this case you're about to read this is what's about Wes law has headnotes and the problem with headnotes is they are often not written by experienced lawyers they are written by people who just got out of law school and frequently are people who just got out of law school and are working for West because they couldn't get a job as a litigator the belief here is that judges are going to do a better job summarizing cases but they know so how to harvest them the one of our one of my colleagues a case text a guy named Pablo arredondo he's an ex litigator and he is the guy who came up with the idea for this technique he put together a regular expression for finding him and you know it's a hit did the job for a proof of concept we had an earlier iteration of the site that had a bunch of these stuck in the margin from this particular technique we have since put together a more refined approach i use a context-free grammar to identify all matching sets of parentheses in the text after that we so see them with citations and see if the contents of the parenthetical match our understanding of what a judicial summary Aria what in these summaries actually looks like there's a few words that always start off the most if it begins with holding there's a really good chance that it's something you're going to want to read we actually sort them based upon the present participle in terms of which ones we think are the most interesting it's a little bit capricious but I think it actually works quite well the other technology we've put together on this is a thing we call key passages and that has to do with the fact that the most quoted sections of the texts are going to be likely to be the ones that you actually are most likely to want to read the and what we find is that people don't quote necessarily the same five or six words but they will quote pieces of the same paragraph and so we do is we find all these inbound quotations and we kind of anneal them together and the end result is these are the paragraphs that are the most significant in this particular case and a example of that go back to oh no so for that you get some really fantastic words from Warren he was a guy who really could write and some of his best writing shows up at the top these are the pieces round before the people quote and they really are some fantastic sections but you know they really also sum up the most important aspects of the case in you know much more detailed language than to get in the summaries so now these two things at the top of the page you already have kind of the most important pieces for understanding this case you might have to read the entire opinion if you are trying to dig a little deeper or more importantly if it's a more obscure case that hasn't been quoted as much as this one but if you're looking in a landmark opinion you are going to get right here everything you might see in a textbook if you were a law student describing the significance of this case which if you are trying to kind of whittle down which pieces of litigate of which pieces of opinion you need to present your brief it's a very valuable tool so that is yeah I think I explained that as far as the actual numbers the first time we ran and I haven't actually updated these since we've kind of tweaks and bugs in it but the first time we ran it we came up with two and a half million summaries over the course of eight million cases and so one and a half million actually got summarized which is a pretty good ratio considering that the graph of citation is kind of a you know big mountain that tails off of the really long tail the most heavily excited cases are procedural things things like Addison Liberty Lobby that's been cited 150,000 times most cases have been cited three or four times often it's not a significant citation in terms of understanding it it's more significant citation in terms of understanding where it went such as say there if you look at the body of the Supreme Court output of it over the last hundred years I think something like eighty percent of all opinions that they've handed down are just simple certiorari denied certiorari denied the you know it's a relative minority of the more they actually heard a case and made that's still significant that's the kind of thing you still want to be able present to the user because it says that the you know the circuit court was final but it it does say that you have this much spread around we actually have done a pretty good job of being able to summarize this stuff without having to use machine summarization techniques which I've gotten a lot better and they least are usually factually accurate but they're not really fun to read so a little bit of a demo this is a word cloud that Pablo mentioned earlier created using the present participles in all of the fountain summaries and you can see holding and finding really do stand out blind share so drag this guy over here it takes a little bigger so I've loaded into the scholar ripple a few of our libraries so one thing you can see we have this judicial somebody's pipeline the chains together all the annotators and if we were to run it so I've fed into it the case that had the very top summary on brownvboard so I'll open that case over here and you can see it and so here is the little summary that we're actually going in looking for and so you can see we have the citation and you can see it's been recognized because the citation has been hyperlinked at this point you have all three parallel sites we go over here again and we run this it will sit here and shop up the case and what you see is all of the debugging output from the grammars that is saying that but at the end of the day it's fine and now we can print all the summaries and so these are the kind of the structures of these wee man occasions you have cavity begin in the end inside of the text and then you have the little pieces of information that we've annotated out of them so for this judicial summary you can see the citation that it was associated with and that's kind of how these will get pulled together but it's great because you can see if one annotation covers another or how else they're related to each other the other thing I was going to right but that's more or less yeah the substance of it will take any questions how big is the data set there are about eight million opinions in overall courts well overall course that we have in our database there are a lot of really minor trial courts and traffic courts that we've not got the opinions of west log would have them because they really do try to be a complete representation of all the law at some point we may try and get that but it's a laborious and expensive to get a lot of that data and it's stuff that most users don't actually care about so it's the kind of thing where we're adding kind of more the long tail more obscure stuff as we grow but there's eight million that cover I think like federal district courts most state courts from like the top three tiers of no state courts say if I stripped all the markup off of it I think probably no more than a terabyte of just actual text we have considerably more in as3 just because we have different representations every time you transform it we stash the artifact and put kind of the fields of the relational database in the s3 key so that there's no need to do a lookup we immediately can just infer what things about have you worked with obstructive summarization at all are just just extractive like I mean I know there's the human created curated pendants are requested a sexy community that can be fried generating such a the abstract aside you know the question was have we tried generating abstractive summaries instead of just extractive and we have not gone that row the assumption was this was really easy to get up and it was going to be the higher quality than our first few passes of abstract it might have been but there are a lot of cases that we will probably have to try techniques like that on just because there's limitation to what you're going to get in a crowdsourced model passages oh so that is i will show you actually an artifact that we get out of that and it will maximize this so if we take this my click on this passage here what it will show is every time somebody quoted a piece of this passage and so now we have 51 times when this particular passage was quoted and I think I just made it go away funny thing is you work on you know the batch processing operation to accompany you find you don't actually use the website that often and then you go to demo when you realize you don't actually know how it works but at any rate this is so you see here we have all 51 cases that quota this pageants you can see the different subsets that they take and the you know I click on this then I can pull up in context where this was quoted and I think that's been a while since i restarted my browser because it's what it's supposed to do is have a little bit more highlighting going on here but sometimes if you you know don't restart your browser for a month stops behaving correctly with everything Sark is actually what does like kind of the running of the analysis so it just takes so it's getting this kind of thing is a to pass process you first find all outgoing edges on every case and you stash that information somewhere part of it goes into the XML so that we can create read hyperlinks other things like that some of it goes into postgres so we have a niche graph and then we do an inbound edge process where we take all the cases and run them through and say okay who whose side of this case and how do they cite it did they quote it do they summarize it and we pull all that data together at once and kind of generate all the pieces of the website there actually isn't really a traditional backhand on the website everything is there's that I don't know if you're familiar with firebase it's a platform as a service over top of MongoDB google bought them a year or two ago that's just used for kind of the mutable user involved data but everything here is actually put in s3 and so it all just kind of gets loaded single page angularjs em anything else all right cool put this back here you