BDSBTB 2015: Koji Sekiguchi, NLP4L: Natural Language Processing Tool for Apache Lucene
Recording: BDSBTB 2015: Koji Sekiguchi, NLP4L: Natural Language Processing Tool for Apache Lucene
you today I'll talk about an introduction to NOP 4l NLP follows in losing is the most successful such software in the world the goal of NLP 4l is to improve the scene users such experience by the assistance of NLP technique there is a famous problem in I our field the trade-off problem between precision and recall they are incompatible I will mainly talk about how we try to solve this problem in our project before going into my talk let me briefly introduce myself my name is Cody cute i'm the founder and CEO of Ramat a such company based in Tokyo we deliver our services regarding loss in and solar especially solar to our customers I'm a committee of apache Lucene and solar so I'm asaji guy as for my experience of scholar I'm not an expert I've been working just own it for just a few month and this is the outline of my talk first off I'll briefly go over NOP 4l then I explain how NOP technique can improve user such experience because some of NLP techniques help improve such a result it is important to understand this in order for you to figure out what's the goal of NLP for L is and why next I'll explain how to count the number of words in the same index it might sound a bit boring to you but it's the first step for calculating probability such as language model or a hidden Markov model I'll talk about these models shortly then i will show you an interesting application transliteration NLP 4l provides it as an example script I'll show you the demo if we have time then I'll share our future plans for this project so what's NM v4l again it stands for NOP to follow scene it's an open source project with Apache License written in scholar our goal is to improve listen users such experience using NLP technique the most distinct characteristics of our project is that it uses the scene index as a co-pastor database instead of using text data directory NOP 4l lets users put their tickets data into loosing index once the data is in words appearing in corpus are settled in losing index trying to account the number of words NLP 4l read loosing index instead of original tickets data there are some other aspects in our project though I will show you transliteration program later as it uses hidden Markov model and when calculating probabilities for hmm it uses loosing index to get the word count in order to better understand later slides let me quickly explain what the scene is the thing is a search engine library written in Java it's apache license and open source with a very active community the same has a variety of functions for information retrieval including indexing surging analyzing parsing dictionaries etc but I want to focus only on this picture the scene makes an inverted index so you can search text document here we have three document that we that we want to search so we inject them to listen and losing in town mix and inverted index from them the index has a list of words appearing in all document each word has posting list there is a list of document numbers of ones that have the corresponding words in them once loosing has the index users can execute qualities for example when you post a koala upon the scene finds it in the index quickly and return the document numbers one and three that includes the keyword next I'll explain how NOP improves such an experience I think it's important to know how NOP technologies can improve users experience when they search by knowing it we understand why NOP for air provides functions that are being implemented now I'm going to illustrate to evolution measures that are very popular in I are filled the precision and the ricoh they are often used in the field of machine learning or NOP so you may be familiar with these measures but let me explain them quickly here we have a rectangle and I'm going to draw a Venn diagram this rectangle represent a universal set this set is a universal set of documents to be searched by the search engine users when the search engine has 1 million documents in its index the size of this rectangle is also one medium here you post your cali let's say kweli x to the search engine when you do that you have your expectation you have a set of documents in your mind that you want the search engine return for you let's call it a target set because it's the target for search engine if the search engine returns a set like the pink circle here the user is happy and satisfied by the result but in reality the such ending cannot fully satisfy you though it could satisfy your expectation partially so here is a blue circle in the rectangle like this it covers a part of the target set as a result the inside of the rectangle will be separated into four sections four sections are composite of the combination of positive negative and through force the result set is considered in positive as the such engine thinks that the set of documents in the below circle satisfy your expectation and the outside of the result set is considered in narrative as a such engine thinks that's a set of documents there are not sorry and there are set of documents that are out of blue circle don't satisfy your expectation they are partially correct and partially in collect from the users point of view they are labeled as true or false let's look at true area true area has two sections to a positive and true negative term means the right answer for example three positive is the area where a set of documents that are returned by the search engine and are wanted by you and to negative is the area where a set of documents that are not returned by the search engine and that are not wanted by you on the other hand falls area can be considered in wrong answers for example false positive is the area where a set of documents that are returned by the search engine but are not wanted by you and false negative is the area where set of documents that are not to return by the search engine but already wanted by you here we have all parameters that are going to be used in our evaluation measures they are precision and Rico they are very popular as you might have heard of them the precision is a major use when you want to represent the placing of the search engine you are using to calculate the precision you use the first of formula TP / TP plus FP or instead TP device by the size of the result set the balloon circle for example if you make a quali and the search engine returns 100 document but you find only three documents that satisfy your expectation then the precision is three percent under Rico is calculated by the second formula TV / TP plus FN or instead TP / the size of the target set the pink circle for example when you execute the search expecting 10 document but the search engine returns only three document that satisfy your expectation daleko is thirty percent now we got eventual measures good and in terms of these measures we as a user / such engine want a search engine that has high performance one that the both precision and recall are high however it is impossible is an and is unable to be implemented the problem is well known there is a trade-off between precision and Rico it's that if you go after high precision you get a search engine with low rico at the expense of it likewise if you go after hi Rico the search engine gives you hi Rico as the cost of precision why this happens in order to implement a search engine with hi Rico you need to adjust this so that it can produce a larger result set then the result set covers almost entire target set as you can see it makes even smaller and you get high Rico number at the same time you lose precision because it makes FP bigger and bigger FB makes your precision lower and vice versa when you want a search engine with high precision you want the original set to be small consequently you get not only high precision but also low recall that's the reason now we understand why there is a trade-off between these measures you can't have it both ways but as a user you want them both how can you do that NLP techniques can help you here this is a solution for having it both ways that I have thought of the important thing is to implement highly go first but as a result you got low precision as i will show you to solve this problem we have two methods one of them is as a bottom left it uses facet and filter quali in order to gradually improve precision the other method is at the bottom right it uses the ranking tuning technique since ranking tuning is out of focus here I want to focus on these other two boxes in order to improve rico there are several method in lachine and so hola such as using a new glum instead of using morphological narrator using synonym dictionary but here i'll briefly introduce you to transfer iteration as an NOP to that may be used to improve Rico and mp4 l has a tool for transportation I will describe it later now we need to improve the decreased pleasure to do that we can use faucet and filter college but these tools aren't always available I'll explain the reason why in the nephews rights in order to make these techniques available we can use named entity extraction there is one of well-known NOP tasks document classification can be used for the same purpose though and focus on name of the entity extraction in the later slides now let's see how the faucet and politically can be used to improve the decreased precision imagine that you visit ebay website and search for what you want to buy when you don't have any particular watches in your mind first you would enter a keyword watch in the search box and click search bottom then you got many watches displayed on your screen this balloon saco is a set of watches you got it should cover almost all watches you want to check but if it stays as it is you would have a hard time finding watches you like because you have so many unwanted watches in your little list that's where the faucet and filter kweli come in you find faucet links on the left of the screen for example there is a link to filter by gender and you would click mint link as it narrows down the site such a result you get the smaller set the other filter queries can be added again you can click another link for example the link of 100 to 150 that appears in filter by price then you get further smaller results yet now the size of the result set is small enough and is easier to find what you want to see now we see that the faucet and filter quality techniques can contribute to improving the decreased precision to implement this scenario document that users will charge must have a structure like this the document needs to be structured using this slide the document hub price and gender field in order for filter clearly to be implemented unfortunately not all documents in the world are structured for example articles in newspapers don't help the structure you see in this slide so one of NLP techniques can help you here name it entity extractor is the answer to this name the entities are distinguished by colors nae is or two to find and extract name of the entities from articles as each named entities belongs to its class or category we can use any e2 to convert our unstructured documents to structured document and eventually we can apply for set and filter qualities to them NOP 4l has an interface to open an LP to use any e function open NOP works nicely but unfortunately it doesn't help modern files for Japanese to implement them we had targeted named entities for Japanese news articles by ourselves to handle this task we used blood there is an open source software for tequesta annotation in total we annotated several thousands of Japanese sentences I'll explain about counting number of words in losing index it looks trivial but I think it's an important first step for NLP this is a corpus which I will I will use to explain about counting words in the next few slides we have three documents in our corpus this is a scholar program that creates Allison index the data source is here we've seen this in the previous slide this is the name of the low seng index directory this is a schema definition it has only one field tequesta field standard tokenizer and lowercase filter are used in this field this function crater loss in document object from a text open later here and use it to write document then close it after learning the code we've got this lossing index that solves our corpus in it now we can count the number of words NOP 4l provides various functions that count words in losing index this is sum total tongue flick and it counts was alkalines frequency in total in the field you specified this is a method to count the number of unique terms in the specified field this is so total prick and it shows use of count of the world you specified in the specific field there are other code snippets you get word count top tones there are two types holding sort by doc flag and total tongue flick let's get further into calculating probability because we've already know how to count the number of words in the Seng Index we can now calculate probabilities interestingly are used loosens single filter to get probabilities so what is single filter it's a word Engram token filter it is used in combination with tokenizer because it is a token filter this is a sample flowchart fight where we use loosens whitespace tokenizer and single filter follows this we put this sentence then whitespace organizers separates it into words shingle filter then produces what by Graham series of tokens we will use single filter later meanwhile let's look at language model language model is often used to represent the fluency of language for example in machine translation the translator uses language model in order to the most natural sentence from several candidates this is a simple simple formula it represent conditional probability of world up or even that word and has occurred because an apple is more fluent than an chocolate the probability of left side member is large enough than the probability of right side member as angular model is most widely used for calculating language model let's see how we do it is simple here to calculate the probability of what Apple even that word and has occurred result of counting the number of the sleeves of the of an apple by the number of the word and so if we if we have a field that organizes word by graham & normal field we can calculate language model from losing index now let's see the code for language model the curve looks almost similar to what we've seen previously except the schema this is a schema we have two fields word and word to g the world field is a normal field whereas words to g is what by graham field it uses single filter and all texts in the corpus is indexed then open the losing index in order to calculate probabilities of the language model then the same pattern can be used for part of speech tagging we will use the same copas again but it has part of these tags and here the table is a description I'd like to use the hidden Markov model to solve the power of speech tagging task this is the formula for hmm what we want to do is on left side we have 30 words here these are given then we want to know the most likely series of part of speech tags as it's difficult to solve this directory by applying an approximation of hmmm on the right side we have two conditional probabilities here and we can calculate these conditional probabilities by counting the number of words or even part of speech tags in the Seng Index as we've seen in the previous slide and if we use our corpus for hmmmm training we get this diagram numbers in this slide are probabilities these probabilities can be calculated by applying the same method here is a corner for hmmm we use targeted corpus here we use hmm model index in order to create losing index that is used to calculate hmm now our corpus is indexed here then we get hmm model and we get hmm toggle to solve part of speech tagging problem here we try to do post tagging for the unknown sentence let's see an inter single application transliteration that NLP 4l provides transportation is the process of transcribing letters or words from one helluva bed to another one to facilitate comprehension and pronunciation for non-native speakers here is an example of transliteration between English and Japanese all Japanese was here originated from English was so they are called long words usually Japanese long words are written in katakana in order to simulate the pronunciation of original words we Japanese often news english words in the document but use katakana for queries or vice versa let's see a specific example here the user such and english word mouse but you got mass in japanese highlighted in such a result what we've learned here is that if we have a list of english katakana word pairs that are originated from english words it helps improve rico for a such engine that's nice and what nicer is that NOP for l has a training data for transportation between english and japanese and in addition to that transliteration program NLP for l has this kind of training data that is a list of english japanese longwall Spears I've gathered them from the internet then I applied from them for alignment this aligned error can be used for training hmmm there is an example scripting NLP for l for transportation here it is it can be executed by loading like this this sample script learns the aligned data and create a model throw in your katana world and the script predicts the most likely alphabetical world and returns it to you this table represents the input katakana world and its predicted English word as you would see the return tunings are nothing more than prediction so some of them are incorrect there is we cannot use them as they are in order to improve Rico so we need a solution this is a solution for the problem this is a system that close web pages on the internet to collect parallel katakana and english alphabet strings it assumes that the katakana world that is close to english was is a long word of the English world to make sure is the long word it passes katakana world into transliteration program then gets the predicted English word from transliteration problem since the return the string is nothing Muslim a prediction it may be correct or incorrect but it should be similar to the real English word here the edit distance comes into play if the edit distance between the collected and predicted English word is less than a certain threshold the problem towards the collected pair of katakana and English words to the dictionary this dictionary is called synonym dictionary and it can be used in Apache Solr in order to increase Rico I got more than 1,800 soon in records from Japanese Wikipedia this concludes the introduction of current NOP for l2 though we have been providing consulting services using NOP 4l the number of customers we have a bit short of what we hope to fall and I should say the business is not quite successful so we decided to provide a framework where various NLP 4l tools can work on NOP 4l framework improves such experience of losing based such systems such as solar and elastic search we provide not only a framework but also the tools which I introduced today as the reference implement we are also planning to provide the transportation copas and in land any model files as well we use NLP and machine learning techniques to output models dictionaries and lawson indexes because NLP and machine learning are not perfect we provide agree that enables you to personally examine output dictionaries this is a volver view of framework the blue flame is a framework while the items inside it functions provided as plug-in there are also are several functions that I didn't have a chance to introduce to you today or that we have not implemented as a way yet moving forward we will define blue allow interfaces and provide implementation that conforms to that interfaces NLP 4l flame work out to this data including word and phrase information that are extracted from sources including copas such engine dictionaries the ultimate dictionaries are such a water suggestion did you mean such synonym dictionary user dictionary for morphological analyzer and key word attachment dictionary in some cases external muscle learning tools such as classification or youth because NOP and machine learning are not perfect we provide agree that enables you to personally examine them you can deploy these dictionaries to solar and elastic search from the gree now let me explain the key word attachment dictionary the key word attachment dictionary is additionally that holds parallel document ID and the keyword list that you want to attach to that document as each items you will refer to this dictionary when you are creating Allison index to attach any keywords to loss in document you can then increase the boost value for that field the key word attachment is considered as a as a general implementation approach for functions including learning to rank personal research named NT extraction and document classification let's move on and let me explain learning to rank and personalized search we usually use cosine similarity between quali vector and each document vector but there may be times we want to use learning to rank line to rank is a task that machine learns data including access logs to obtain an appropriate search ranking please refer to line to rank on Wikipedia for the detail let's think about realizing this with listen program learned from access log and other sources that the scholar of document d should be larger than the score that the scene normally calculates for qwali q and as you can see in this diagram it performs keyword attached cual eq in the document d increasing the boost value of this field makes the longing of this document go up from the next search personal research means optimizing the search ranking for qwali q on the charger by such a basis likewise program learned from access log and other sources that the scroll of document d should be larger than the score that losing normally calculate for query q by user you since you cannot pass you to the school function parameter as loss in restricts link so you have to combine quali Q and the user you to create QW now this test can be realized by Q at attachment as well but you have to devise ways such as limiting the data to high order queries or dividing lezyne field depending on users because the number of clearly user combination can be enormous okay that's it we are looking for engineers who are willing to work with us on NOP 4l contact us if we you are interested in our project and like the earlier including information retrieval NOP and machine learning thank you awesome thank you 30 you