Hidden in plain sight: Using law to summ...
Summarization remains one of the most conspicuously unsolved problems in NLP today. There is plenty of active research, and the best available is reasonably capable of capturing the information contained in a body of text, but the output is often clumsy compared to what might be written by a human. More importantly for those of us working in legal informatics, the law is a notoriously conservative profession. Lawyers will justifiably be more comfortable relying on summaries written by fellow members of the bar. Fortunately, judges provide high quality, detailed summarization of the cases they cite to in their opinions, and a rich body of this data exists throughout the body of historical case law. We present a technique for enriching the display of judicial opinions with high quality summary data extracted from subsequent opinions, using a variety of state of the art open-source software tools. FOSS tools used include Antlr, for recognition of deterministic sequences using formal grammars, and Apache UIMA for construction of multi-layered indexes of recognized entities, such that their various juxtapositions can be used for further inference.