Scale By The Bay 2019: Rajesh Muppalla, Lessons Learnt Building Domain Specific NLP Pipelines
So I'm going to talk about domain-specific NLP pipelines. Yeah, quickly about myself. So I'm a senior director of engineering at Avlara. So I take care of the, you know, AI and ML platforms. And before this, I was a co-founder at Indics. And my company was acquired, you know, earlier this year by Avlara. And prior to that, I was a tech lead at, you know, ThoughtWorks, working on a product called GoCD, which is an open source CI and CD tool. And, you know, this is my third time speaking at scale, by the way
An interesting story of, you know, how I got associated with this conference. So, you know, I think exactly about five years back, I met Alexei on a plane. I think both of us were coming back from Portland after attending a conference. And, you know, I just spoke about some of the stuff that we were doing back in Indics, the data pipelines and machine learning. And he said, you know, I run this conference, you should come and talk. So that's why I was invited to be a panelist at, on data pipelines. And then a couple of years back, I spoke about, you know, how to apply the current delivery principles in machine learning. Basically, you know, how do you take models to production
And my team is based out of Chennai. And yes, I'm traveling all the way from India for the conference. The focus of this talk, so we've been working on, my team has been working on NLP problems for the last, you know, seven odd years. Across two domains now. Of course, at Indics, which was an e-commerce company. And the NLP stack there evolved over a period of about seven years. And this year, you know, we got acquired by Avlara, which is a tax company, which is in the tax domain. And we, of course, you know, applied some of the learnings from our e-commerce domain into the tax domain
And we also have figured out, you know, some of the new techniques that have emerged over the last couple of years. And this talk is all about, you know, sharing those lessons. Okay, quickly, I want to talk about, you know, Indics. So our mission was to be the Google Maps of products. What do I mean by that? You know, you've, if you want to get products, of course, you know, there's an Amazon and eBay. But there's not a single source that can give you all the products. So we wanted to be that source. And we wanted to, you know, create a catalog
At this moment, we have a catalog of about three million products that will enable customers to, you know, build product-aware applications. And the way that we built this catalog was primarily by crawling. So we crawled about, you know, 5,000 retail websites across the world. The next thing was parsing. So this involved, now each website that, retailer that you crawl would have their own page structure. So, so we had to parse those pages to extract the semi-structured information. So this involved computer vision. The next thing is, you know, classifying these products
Every product that we, you know, that we extract. The taxonomy that, that these retailers would have would be very different. So we had to, you know, categorize this, that into a single taxonomy that, that we had of about 5,000 leaf nodes. And then one of the most interesting problems here is matching. The same product being sold on multiple stores need to be, you know, gathered together so that, you know, it enables insights, insights for downstream systems. So that, that process is called matching. And doing this at scale across 3 billion was, was a hard challenge for us. And then finally, this entire database is available for customers as data as a service
So they could use, you know, through search or API to, or, or even through feeds to access this, this data. Now let's look at, you know, how the NLP stack looks like for an e-commerce domain. And this is, of course the stack has evolved over a period of seven years, but this is the latest snapshot. So on the, on the bottom you see, you know, raw data. So there is, there is, there is data. There is raw data which is unlabeled and training data that is labeled. And, and one of the things at index was that, you know, we, you know, over a period of seven years, we build a very good corpus of label training data. Which is not, you know, something that, that is available for, you know, for, for any company that's starting with NLP
Um, the next step is parsing. So, you know, you have, you have documents. So in our case we have, we had HTML pages. So we had to write parsers to extract, uh, uh, uh, this, uh, the information from the HTML pages. Um, the next layer is of preprocessors. So, you know, tokenizers, you know, uh, uh, lemmatizers, PR tagging and language detection. These are some, you know, fundamental preprocessing, uh, constructs. Again, I mean, this is pretty common, uh, you know, in, in an NLP stack
And the next one, which is one of the most important, uh, you know, building blocks of, of any, NLP, NLP stacks, involves, you know, embeddings, language models, and knowledge graphs. So embeddings are nothing but, uh, you know, vector representation of, of, of your text data. So we used character and word embeddings, you know, uh, including, uh, examples like, you know, word to work and fast text. Uh, we also use language models. You know, these are, uh, n-gram based statistical language models. Uh, you know, to solve some of our search problems. And then, most recently, we also built a knowledge graph. Uh, this is to augment, uh, you know, uh, our models by, you know, giving it external data
So we had, uh, about 100 million nodes and, uh, about a billion plus edges. Um, and, and the next layer is of algorithms. So we have, uh, uh, I mean, there's classification, NER. Uh, uh, then there is document similarity, you know, auto complete inquiry expansion for, uh, uh, for, for the search use cases. And then, of course, the use cases, which is part of classification, you know, at word extraction, matching, and search. And then this, on the, on the right side is the technology stack. Uh, so we, I mean, for, for our data pipelines, you know, we have, we, we primarily use Spark. Um, and we are, we are a Scala company, but over the last couple of years, you know, we've started using Python more because the ecosystem of libraries is, is much better in Python as far as NLP is concerned
Uh, for pre-processing, you know, Spacey is our go-to library. Uh, D-graph is a library that we use for knowledge graph. And, uh, as I said, you know, fast text is, is, uh, I mean, what we use for classification. And Spark MLlib for some of our NER work. And then for, uh, model versioning, um, and, uh, tracking our experiments, we use MLflow. Um, so one of the things I want to focus on is embeddings. I mean, as I said, you know, it's a, it's a fundamental building block of, of our NLP stack. So, so, so what are embeddings? Uh, so, but the definition of embeddings is it's a dense representation of word vectors learned from a large unlabeled corpus
So why do you need, why do you need embeddings? Uh, unfortunately, you know, machines do not understand text. They only understand, uh, numerical data. Uh, in case of images, you know, you can, of course, take the RGB data. Uh, you know, you get, you get, you get, uh, the pixel values. And, you know, that, and that you can convert into numbers. Right? But in text, you don't, you don't have that facility. Right? So you need, you need a way to convert text into numbers. And, and this is part of feature engineering
And one of the things, you know, we learned is that, uh, uh, I mean, the quality of embeddings matters a lot. So the quality of your models, you know, depends a lot on the quality of the embeddings. And most of the times, you know, these, these embeddings are learned. You can also, of course, create them by hand. But most of the times, you know, these are learned. Um, and these embeddings, you know, capture a notion of similarity. Because similarity is an important aspect for, for downstream NLP, NLP tasks. And, word embeddings were popularized, popularized by word2vec by, uh, a person called Mikhalov in 2013
And, glove and fastex are, are, are other implementations of, uh, of word embeddings. Now, it turns out that, you know, embeddings have some useful properties. Uh, uh, so they capture certain relationships between words. Uh, for example, the vector representation of, uh, uh, I mean, of, of king and, uh, I mean, the vector representation of queen, right? I mean, they have, they have a similar, uh, let's say, difference between, uh, the same as, let's say, man and woman. Uh, this gives you some interesting, uh, interesting things like analogies. For example, if I take the vector for king and subtract the vector for man and then add the vector for queen, I would get the vector for woman. Uh, so, yes. Um, and then you all, I mean, similarly, you know, uh, these, uh, these embeddings also capture verb tense as well as, you know, country and capital relationships
Now, how do you train these embeddings? So, this is an example of, uh, word to wack skip, skip gram model. So, essentially, uh, I mean, word to wack, you know, there are two, uh, models. You know, there is a C bar, which is, stands for Continuous Bags of Words, where, uh, you know, the training objective is, you know, given, uh, uh, set of context words, you have to predict the middle word, the, the central word. But in case of skip, skip gram, it's the other way around. Given a context word, you have to, you have to predict the, uh, the, uh, the context words. So, so, so, the way that you do training is using, using a neural net here. So, the neural net has one input layer. It has, uh, you know, just one, one hidden layer
And then, and then, uh, uh, output layer of, uh, uh, you know, which is a softmax classifier. Uh, the, the input layer is a one-hot, uh, uh, one-hot encoded vector. Uh, so, so the idea here is, you know, for every, I mean, you, you start with, uh, with an entire, you know, uh, you know, unlabeled corpus. You, for every word, you, you, uh, you know, do a one-hot encoding for it. And then, you pass, you pass it through the network. And the idea is that at the end of it, uh, I mean, uh, the, uh, the network should output the words, you know, that, that, that, that are closer to the context word. And then, if they don't, you know, you, you actually go back and, uh, you know, update the weights accordingly. So, so now, uh, so now what happens is, I mean, you, you need to do this for every word
And you might also have to do multiple adaptation over the entire dataset. So at the end of it, what happens is you get, uh, you get the hidden layer, which, which now is, is, is your, is your weight, weight matrix. And then, uh, and this weight matrix becomes your, your, your, your word embeddings. Now, now, how, how is it, how is it that this weight, weight matrix becomes a word embedding? Uh, so if you take any word, uh, any word and then do one, one, one hot encoding of it and multiply it with the, with the weight matrix, what you get back is just that one row, which is the word embedding for that row. Now, let's look at some examples of, you know, where we've used the word embeddings at, uh, at index. Um, so one example is, uh, the matching, uh, matching problem I spoke about, which is, you know, two products, uh, from different retailers, are, are they, are they same or not? So, so what happens is, I mean, you know, you have, uh, these two products, uh, these two titles at the bottom, lib, lib, um, so, so this one, I mean, sorry. Yeah. Okay
Yeah. So, so, so, so, so the title there is, I mean, on the right hand side is Burt Bees, lib, um, I mean, this is one of the products. And the other product there is, uh, uh, lib bomb by Burt Bees. So the first thing that, that happens is, you know, you, you need to get the word embeddings for this. So, I mean, so, so that happens here. And similarly on, on, on, on, for this other product as well. Um, the next step is, you know, you would, you would average these, these, these word embeddings. I mean, you could also, of course, do concatenation
But for us, you know, averaging of these word vectors, you know, help. So now this becomes, uh, a document vector. And now the two document vectors on both sides, you do a cosine similarity. Uh, and then based on a threshold, you decide whether these products are same or not. Um, so, so another problem is, uh, the product classification in the e-commerce domain. Uh, so, so these are four products available across four retailers. So, if, if you look at, you know, it's the same product, but there is subtle variation in the title as well as the breadcrumb. Um, but then, you know, we have to categorize all of these into a single, uh, category taxonomy
Uh, in this case, it's lip care. So this is, uh, so, so this is a product classification problem. Now let's see, you know, how did we solve this, uh, uh, using word embeddings and fast text? So, very similar to what, what we saw in the document similarity problem. Uh, you, you have, I mean, this, assume that, you know, this network has already been trained and we've got the word embeddings and, and, and, and the weight matrix. So, you first, uh, you know, get the word embeddings for, uh, each of the words. Uh, and then you concatenate them to get your, uh, sorry, in this case, averaging. You average them to get your, uh, uh, your, uh, your document vector. And then it passes through the hidden layer and the output softmax layer
And out comes your, your category probabilities. So in this case, you know, so you can see that lip care is, is, is, is a higher probability. So, so, so, so this is where, you know, we've used, uh, uh, uh, you know, word embeddings and fast text for classification. Yeah. So, so that was index. Uh, so now let me talk about Avlara. Um, so Avlara is a, is a tax compliance, uh, company. And, uh, you know, our mission is to be part of every transaction in the world
Um, so, so the way, so the way that, uh, that our, our platform works is, you know, we have, on one hand, you know, we have, we have countries about, uh, we support about 203 countries. Um, and, uh, and, and, and, and each of these governments, you know, if you see, for example, if you take US as an example, uh, you have, you know, state governments having department of our new website, where they publish, uh, you know, tax rates and, uh, and, and taxability rules. So, so, so my team actually, you know, goes and crawls these websites and tries to extract, uh, the tax rates and rules information. Uh, the next part is the actual, uh, the core of this entire platform called our tax engine. Um, so this is where the tax determination happens. So there are four, four aspects of, of any tax determination. First is what, uh, where, who, and when. So what involves, you know, the product being, uh, the product of the service being sold
Uh, for example, you know, there'll be some states where, uh, uh, let's say, you know, diapers are, uh, uh, are exempt because they're essential. But cigarettes or alcohol might have a higher, uh, uh, tax rate because, you know, they're luxurious, right? So that's, so identifying what product or service is being, being sold is, is important. The next thing is, uh, who, uh, so, so, so depending upon, you know, uh, you know, who's, uh, so, so let, let's say if one of the customers, uh, you know, has, has an exemption. For example, you know, their agriculture or in the military. So they would have some exemptions. So, so based on that, uh, you know, that the tax rate for them will be different. The next thing is where, uh, so every jurisdiction, every state has their own tax rates, right? So, um, so let's say if, if, if I'm, I'm somebody, you know, who's selling to a customer in, in another state. So I might have to pay tax in California as well as in that state as well
So, so there are, you know, subtle rules around that. Uh, the final thing is when, um, so, and these tax rules might change, you know, monthly, weekly, uh, not, not weekly, but, you know, monthly, quarterly or, or yearly. So depending upon, you know, when the sale was made, uh, what is the tax rate that's applicable needs to be applied. And there are also these special things around, uh, you know, uh, sales tax holidays. For example, back to school, uh, holidays, you know, where, uh, the tax rate is different. Uh, you, you, I mean, uh, so any stationary items you buy for school, you know, there's no, there's no tax rate. So, so, so again, you know, when, when, when part becomes important there. Um, and then we also, you know, have a lot of connection, uh, connections to, uh, you know, ERP systems and, and other, you know, POS systems for us to, you know, get, uh, get access to transactions
Uh, so that, you know, we can calculate, we can, we can get the transaction and calculate the, uh, tax at real time. Uh, we also take care of returns for customers. So here again, there are issues around, you know, every, uh, jurisdiction having their own, uh, you know, return form. So how do we make sure that, you know, we automatically file returns on behalf of our customers? Now, now let's look at, uh, you know, some of the problems that, uh, some of the NLP problems that we're solving at, uh, at Avlara. Uh, so look at this example of, you know, tax rules in New York. So, so there are taxes for specific types of clothing. Um, so you can see that, uh, you know, on the left hand side there are, you know, examples of taxable, uh, clothing and on, on the right hand side you have non-taxable clothing. Um, so it's not just, uh, you know, all caps are taxable or all caps are non-taxable
Sh-shower caps are taxable but swimming caps are not. Right? Similarly, you know, walking boots are taxable but hiking boots are not. So, so you can see that, you know, there is, there's a lot of complexity in, uh, in, in, in the taxation and, and a lot of subtleties involved. You know, I can't just classify a product at a, at a higher level of clothing. I have to specific, or not even caps. I have to classify a product at, at the granular level of either shower caps or swimming caps. So now, you know, let's, let's understand the whole, uh, classification problem, uh, you know, in the, uh, in the tax agreement with an example. So, so the same birds, bees, lip balm example if you take, uh, and, and assume that, you know, we, we, we got this as part of a transaction, uh, in, in our tax engine
The first thing it needs to do is it needs to classify and, and assign it a tax code. So in this case, the tax code assigned is skin care. Right? And, and, uh, you know, these are, these are, this is a very, uh, Avlara specific taxonomy. We have about 2000 such codes. Uh, and, uh, and the idea, and the reason we have this tax code is so that, you know, we can associate rules with, rules with these tax codes. So on the right hand side, you can see all the rules associated. So for example, in this case, uh, for, for, for the tax code SK0001 in New York, the rate is 8%. So in addition to, uh, US based, uh, tax codes, we also support, uh, cross country, uh, sorry, cross border
Uh, now let's say if I ship from one country to, country to another. So, so there is this, uh, you know, uh, tax code called, uh, harmonized system tax codes. You know, these are tariff codes. Uh, these are about 8000 of these and, and every country has their own, own, own version. So we have to solve this problem, not, not just for, uh, US states, but also for, uh, you know, uh, countries. Uh, you know, uh, cross border shipments. Now, if you look at this classification problem, you know, you would say that, okay, you've solved a similar problem in the e-commerce domain. So it should be, it should be easy, right? Uh, however, you know, there are some challenges
So the first thing is your class, classification taxonomy is different. Uh, you know, when we were solving the problem at index, we had about 5000, uh, leaf nodes. Uh, but, but here, you know, the, the, there are about 2000 leaf nodes and the labels are very different. So the, so what, so the category labels are, are, are different in this case. Um, the, the other thing is, I mean, at, at, at index, you know, one of the things we realized, and we were spoiled with, a lot of label data, which over, which we built over a period of time. But here, you know, we didn't have a lot of, uh, label data for, for us to train. So, which is, which is called the small data or the low data problem. And the vocab, vocabulary, uh, of products is also different, uh, across the domains
Uh, so what we noticed is that, uh, the data, uh, in the tax domain, the transaction data that we get was a lot more noisier. And, uh, you know, a lot of abbreviations were being used. So, so, so the vocab wasn't like completely matching. Um, but then, you know, we do have, we did have word embeddings from the product domain. So, so what we were able to do was, you know, we were able to create a decent baseline by, by using them. Uh, and the question is, you know, can we do better? Um, and the answer is, you know, yes, yes, we can. Uh, so these, these are some of the techniques that we've applied, uh, you know, uh, recently. Uh, and, and I'll talk about each one of those
Uh, transfer learning is something that's already in production. The other two, you know, we are still experimenting. Uh, you know, it's, it's, it's been some time now. So let's first talk about, uh, transfer learning. You know, before going into the details, uh, of, of transfer learning, uh, you know, I, I, I just want to talk about, uh, an analogy here. Uh, so as humans, you know, when we have to learn a new task, uh, uh, which is, you know, might be unrelated or related to some of the, some of the tasks that we've already known. Uh, so we don't actually start from scratch, right? We, you know, we, we pick up, uh, you know, what we've learned from earlier tasks and then, uh, and then you apply that in a, in a, in a newer task. You know, another example is, uh, you know, I, I, I told that, you know, I spoke about, uh, you know, content delivery for machine, machine learning, uh, topic, right? Uh, so I had, I had some background in content delivery and machine learning was new for me
So, but when I, when I, uh, was started working machine learning projects, I saw that a lot of pain points that we were facing were very similar to what we were facing in, in software. And then, you know, CD solved some of those problems. So I tried applying those. So, so, so the analogy here is, you know, you don't always start from scratch when you learn something new. And, and that is what, uh, you know, transfer learning is all about. So if you look at traditional ML, you know, if you need to solve a problem, you have a data set. You know, you, uh, you know, you, you build a model for it, right? And you need, you have, you have another problem. So, you know, uh, for which, for which you need to solve
So you, uh, so there's a separate data set for it. So you don't, uh, I mean, you, you build another model for it. You don't reuse between the two. But in transfer learning, so what happens is, uh, uh, the knowledge that is learned from one of the, one of the systems, you know, can be reused in the, in the other system. So that's the essence of transfer learning. And, uh, you know, let's talk about the history of transfer learning. So transfer learning is not new, per se, you know, it's new for NLP, but it's been there for a while, uh, in the computer vision communities. This is, computer vision is nothing but, you know, dealing with images
So in 2012, uh, you know, there was an ImageNet competition that, uh, where, you know, people had to classify a set of images. And, uh, you know, there was a deep neural net, uh, using, you know, com nets and, and transfer learning that was built called AlexNet, which beat, uh, the competition by about 41%. And it had used transfer learning. So that's when, you know, people started noticing it. And then, uh, you know, there were, there were, there were a lot more models like that built in, in, in the, in the computer vision world. But NLP, you know, it took a while. Uh, and last year, 2018 was a, what people call a watershed movement. Uh, where, uh, lot of people using transfer learning, uh, in, uh, a lot of researchers started using transfer learning, using pre-trained language models
And these, uh, you know, architectures came up like ELMO, ULMFIT, BERT, GPT and GPT2. I, you know, we will talk about some of them, uh, a little later. Um, and, and this is, uh, a quote by one of, uh, you know, one of the researchers from DeepMind. So he says that it only seems to be a question of time until, you know, pre-trained word embeddings will be dethroned and replaced by pre-trained language models in the toolbox of every NLP, uh, practitioner. You know, this is something he said, you know, last year. And we're already seeing, uh, uh, you know, this, this, this happening. So to understand this better, I mean, let's actually take, take one example. So we'll take ULMFIT, you know, uh, it's, uh, uh, it's one of the pre-trained, uh, language models
Um, so the first, you know, step is you pre-trained, you pre-trained it on a, on a source data set. So the assumption here is that, you know, you have a large data set. In this case, you know, we have the e-commerce product data which is large, you know, which is, you know, millions of, uh, uh, uh, records. So what we can do is, you know, we can actually pre-trained, you know, we can create a pre-trained language model on top of it. Um, so, and, and this is architecture used, uh, so this is called AWD LSTM. Uh, it's a, it's a, you know, deep neural network architecture where the first layer is an, is an embedding layer. Uh, and then, and then you have a hidden layer which is a three-stack, you know, by LSTM. And then you have a softmax layer
And here, what you do is, you know, you pass this data set and then you train this network. And the objective function here is to predict the next word. And, and, and, you know, and, and given the size of the data, you know, this might take any, anywhere from, let's say, you know, a few days to weeks. Yeah, and, and, and, uh, uh, and this is called a pre-trained, uh, language model. Um, and these are some examples of, uh, pre-trained language models. So, uh, they differ in architecture. So, so we were looking at ULM, ULM fit, which, whose architecture is AWD LSTM. Uh, there's also ELMO
Um, and then, you know, Transformers is a pretty popular architecture and examples of, uh, language models within that are GPT-BERT and, uh, GPT-2. Uh, they also differ, uh, in terms of what is the objective function that, that gets used, uh, in case of, uh, you know, BERT, in addition, in addition to the, predicting the next word, you know, they predict the mask word as well as the next sentence as well. Uh, and then the number of, you know, parameters that go in, uh, into their, uh, I mean, into the architecture is also, is also very different. I mean, for example, again, in BERT, there were 110, uh, 10 million parameters. Now, in the first step, we, we, we have pre-trained the language model on, on a large, uh, uh, uh, you know, dataset. Now, the next thing you do is, you know, you fine tune it on the domain dataset. So, in this case, uh, the domain that we're talking about is the tax domain. So, so what you do is, you know, you, you take that dataset, uh, which is, which is right now unlabeled
You don't care about any labels here. And then from the pre-trained language model, you take the embedding matrix, which is, which is the, the bottom layer, and the hidden layer, which is the weight, weight, weight matrix. Uh, and then you, you, you train here. the objective function, again, is same as what was in the pre-trained language model. You, you predict the next word. So the difference here is that the, the, the dataset is, is, is different. You know, it's your, you know, specific, you know, domain-specific dataset. Uh, and then you, uh, I mean, there's a, you know, there are, there are certain tricks that you apply, uh, you know, during the fine tuning
Uh, so one of the first, uh, I mean, uh, one trick is what is called freezing. So, your, uh, pre-trained language model has learned, you know, a lot of, uh, nuances, uh, you know, because, because just to predict, predict the next word, right? Uh, uh, you need to be able to understand, you know, uh, you know, how, how sentences are formed, the semantic and the syntactic nature of, uh, uh, you know, nature of the language. So you don't want to lose that. you know, while you're, while you're fine tuning for your specific domain. So what you do there is, you, you freeze your, uh, uh, I mean, your hidden layer. And then you, you know, and this is only for a few epochs, you train the rest, uh, the rest of the network. And once you're done, I mean, once you're done training, you then, you know, gradually keep, keep, unfreezing, you know, layer by layer. So this makes sure that, you know, whatever, uh, the language model has, has learned, the pre-trained language model has learned, it doesn't lose while the fine tuning is happening
So another technique is called, uh, you know, slanted learning rates. So where, uh, the idea is, uh, uh, you know, I mean, learning rates is, you know, how, is, is, uh, is, you know, how you update the weights, you know, how fast you, you, I mean, you, you update the weights. Um, so here again, you know, you should, uh, I mean, uh, what we do is, you know, we use a, you know, stop, I mean, we, we use a steep learning rate initially so that, you know, we, we get near the parameter boundaries. And then after that, you know, gradually, uh, change the, you know, change the learning rate so that, you know, we are just doing fine tuning. So now, now you have a fine tuned language model on your specific domain. The last step is, you know, you, you build the classifier. I mean, you, you, uh, you know, you, uh, you actually then, uh, do the final task of classifying for your, domain. So in this case, you know, we, we need to do a classification
So the data set in this case would be, uh, the, the tax domain, the label data set, which is like, much smaller than, uh, you know, than, than what, than the previous two data sets. Um, and here, you know, what you do is you, you remove the, you know, the top most, the softmax layer that was, there in the, the language model. Uh, and then apply your own, uh, you know, layer which is specific for classification. It need not be just a deep, uh, you know, deep layer. It could also be even a shallow layer like an SVM, for example. And here the objective function is, is, is what would be of a typical classification task. In this case, it's cross entropy. Uh, so classification was one example
We've also used, uh, you know, transfer learning in this, in this example, which is, as I said, you know, we, we, we extract tax rate changes from, uh, uh, uh, tax notices, uh, automatically. So in this case, you know, you, you can see, this is a notice for, uh, tax rate change for, uh, a state of South Carolina. And on the right hand side is the structured information we need to extract. So this falls under the, uh, the problem, NLP problem of, uh, you know, semantic role labeling or slot filling. Uh, and we also need to co-reference resolution when, uh, there are more than, you know, one sentence in picture. Here we used, uh, ELMO. Uh, and in this case, you know, we didn't do any fine tuning. We just used whatever pre-tale language model we got, uh, to, to, to solve the problem
Uh, so one of the questions, you know, why does transfer learning work? Uh, so, so many NLP tasks share common knowledge about language. So, you know, the, so when you look at language, you know, the linguist representation and, you know, there are a lot of structural similarities in, in, in the language. So, so, so that's why, I mean, you know, if, what, what is learned from one, from a large corpus is still applicable in, in, in another, uh, another corpus. And of course, you know, tasks can inform each other about the syntax and semantics. Um, and, and it's also important because, you know, label data is rare, but unlabeled data is, is abundant in every domain. So, so, so, so I mean, if, if you have a lot of abundant data, I mean, there is, there's no reason for you to not use, you know, transfer learning. Now the question is, okay, I mean, you know, ultimately both are, you know, represent, vector representations of, of your textual data, right? So why is, so what is the key difference between word emitting and PTN language models? So when you look at word emitting, it was just a single layer in, in, in your network, right? But, uh, as far as PTN language models are concerned, it, it, it's, it's a full, full, you know, network that, uh, you know, uh, uh, that is there. So, because of that, what happens is, you know, it, it helps you capture context
So in case of word emitting, you know, if you look at this example, uh, there are two words, I mean, two sentences, but, uh, uh, the meaning of apple is actually, uh, actually different in, in both of them. The word emitting, I mean, they'll, the, the word vector for apple will be same, but in case of, uh, pre-trained language models, you know, once, uh, you know, once this, uh, this sentence passes through the, uh, through the network, because of the, context, you know, it is, it, it, it, the vectors for apple and, uh, apple in both of these will be different. Uh, there are a couple of other techniques, you know, weak supervision is an, is, is an example, is a technique that we are, you know, currently experimenting with. It's also called data programming. Again, the goal here is to increase, uh, label data. So, the, some of the libraries that, uh, you know, we use are Snolker and Snuba, Snuba. Uh, I, I, I mean, this, uh, the, the, the, the, way that you, you know, go ahead and do this is, you know, you take your label, uh, label data, you know, you, you look at it manually and, you know, figure out some domain heuristics or, or, or, labeling functions. Uh, actually, Snuba can do this automatically and Snorkelly have to do this, uh, you know, manually
And then, you know, based on these label functions, uh, a model is generated, uh, that can emit, emit, you know, probabilistic labels on your, on your data. You take this model and then, you know, run it on your entire unlabeled data set. Again, you know, it will generate, generate product, uh, you know, probabilistic labels there. And then, you now have a good enough, you know, training data set which is also large. So, that's, so, this is one example of, you know, how you can increase the training data. Uh, another example is, uh, you know, data augmentation and, uh, uh, and synthesis. Uh, so, I mean, this is one technique where we used, you know, we translated from one language to another. Uh, and, and then translated back again
So, you, you are able to find some subtleties between the two. Yeah, I mean, and, and this is the final, uh, you know, how the, uh, NLP stack for tax domain looks like. You can see that, you know, pre-trained language models, models comes, uh, comes into picture here. And in terms of technology, I mean, uh, you know, uh, so PyTorch, LNNNP and FastAM, these are libraries that we are using, uh, right now. So, so conclude, uh, you know, NLP pipelines need to be domain specific, uh, uh, so there are libraries, infrastructure and techniques that you could reuse across domains. Uh, but having good quality and, and label domain specific data is of, uh, you know, utmost importance. Um, and, and if you look at specifically in domain specific data, you know, there is, if you have large, large label data, I mean, a lot of things will, a lot of algorithms will, will work for you out of the box. You don't have to worry about, you know, some of the techniques that I showed up, showed you
Um, and you should always try to use, you know, unlabeled data from your domain, uh, uh, to your advantage. And in case of, you know, if, uh, label data is, is small or low for you, so you can use, uh, pre-trained language models, transfer learning using pre-trained language models, which I demonstrated. And of course, you know, techniques like weak supervision and data augmentation will also help. Okay, yeah, thank you.