Bay Area AI: Alyssa Wisdom, Topic Modeling with LDA
Recording: Bay Area AI: Alyssa Wisdom, Topic Modeling with LDA
[Music] great so welcome everyone today I will be talking about topic modeling using late and derelict allocation my name is ELISA Lissa and give you a little blurb about me before I dive into the talk today so apparently I am a product analyst at square and my current role and responsible for using data science and bears in eluc tools to help inform the products and how to move I to move and where to move essentially yeah there's previously I worked in growth analytics at a startup called quarter ahead and I was there's a lot of hacking work and helping with wrote them attention for that status report and before that I was doing merger and acquisition work in Singapore for a few years so turn universe background I have our bachelor's and a master's in psychology in Campton and if you ever have any questions about anything that I cover to be in the top or you just want to talk to me about all things data because it's kind of cool feel free to email me at a wisdom that's grab calm make it pretty easy yeah all right so for the top today I'm gonna be a little bit of there's a little bit of application I think there is important to cover so I'm going to be talking about just machine learning basic machine learning concepts topic modeling and how it fits into the machine learning framework lain here with allocation which is the major topic modeling method that I'll be diving into today and different model evaluation metrics that you can use to make sure your models doing what you actually want it to do and then for the second half of the talk I'll be showing you different applications so talking about the Python libraries that you can use to do way in their application and showing you a example with some Python code so let's start with assume learning so how many people have seen something along the lines of this before yeah okay so I mean whether you're a consumer or you're working in a company we've all seen something pretty similar right when we unsubscribe from a product this is a free text field which we show users to kind of gauge why they're leaving the product some companies use categories instead of free text and you know these categories are usually based in some type of intuition or maybe past user feedback some of the pros of having set categories for users to choose from is that it's really easy to understand traps right the product manager can just kind of look through it see how many people will show us which reasons and figure that's the top reason but one of the drawbacks of having preset categories is that it's really very confining and they don't they're very slow to evolve with the customer so that's why a lot of companies supply squared use free text fields to do this the upside to it is that you get a lot of tech a lot of rich text qualitative data but the downside is that it can be really difficult to parse so my campaign knew one day and he said hey Alyssa you know looking at the pump as we always do with SAS products looking at how we can either increase acquisition or a decrease churn and he said I wanted to look into retention and we have all of these turn reasons and I don't want to press through all that in Excel slash don't do we know how so can you help me and the first thing that came to mind was okay this probably calls for there's some type of machine learning solution using topic modeling okay so first let's talk about the world of machine learning and how topical topic modeling fits into this framework the way that I've seen machine learning is that there are tons of options and applications at your fingertips for you to use and depending on what kind of data you're working with you really choose the best application that needs to do in these there are two major groups of machine learning applications which is supervised and unsupervised learning this is essentially determined by whether or not your data set does or does not contain label target outfit values in supervised learning you have a set of inputs usually a vector or numerical representation of features of some sort with corresponding target output values it's also known as the Supervisory signal supervised learning algorithms analyze and they learn from a lot of training data and they produce an inferred function which you can use for mapping unseen inputs to their corresponding target values it's a targets are categorized into classes then classification followed if the target is continuous at quantitative then it can be a regression problem in unsupervised learning though you only have a set of inputs with no corresponding target value so essentially your training your machine to learn from the inputs and figure out the late in or the hidden structure and relationship between the various inputs it's really good for exploratory analysis and under and really understanding how things are really then supervised and unsupervised learning have different use cases depending on what problem you want as well so let's say you're a major soccer field fan which I am I'm into what jackie film okay okay we have some subjective no fans here that's great and last thing you wanted to create a model that could predict who would win the men's 100-meter dash right and 2016 Rio Olympics you would probably look at the past five years of World Championship data for 100-meter dash and keep the names of the competitors as the inputs and the binary outcome of the competition which would be win or lose would be the outputs in your to train your model based on the training data you would probably likely learned that in the same Bowl one of the most successful track and field athletes in the history of the sport have a track record of winning and would probably predict for neither win you wouldn't be wrong now sake or let's say that you're Vic Schoen who's a director of head director and head coach of Jamaica's national tracking bill team who helped usher you saying into his professional career if coach Coleman wanted to better understand what three major characteristics might into making an elite athlete to better guide and develop his athletes like insane he could do this with unsupervised learning by looking at the characteristics of major accomplished athletes and let's say in the past ten years or so such as like weight height hours practice a better performance stats etc he then used these stats as inputs for a clustered model which were telling something along the lines that these characteristics indicating natural ability maybe work ethic and mental fortitude are the major pillars of accomplished athletes or le appellants so his ruthless characteristics into three zip that stop the amount of cussing okay so based on this what kind of I guess for everyone I'm not quite sure if everyone is familiar with hopping modeling or not but if you're not familiar what kind of machine learning application do you think topic modeling is based on those examples the people not familiar with it anyway just caught up unsupervised right so talking ball is a is a type of unsupervised machine learning that makes us makes use of clustering and looking freely in variables or Canadian structures in your data set one type of topic modeling that I prefer to use is lame derelict allocation not to be confused with linear discriminant analysis has been standard acronym and I mean everyone who uses Python when you're looking through like scikit-learn for different libraries that you can use um Elvia is a Lda is something that is there don't use linear discriminant analysis make sure you are very mindful of that see a lot of threads where that's gone wrong on a number of occasions so yeah the real Lda as I like to call it is a generative statistical topic model for finding accurate mixtures of topics with the given document set or within a different given document set and did you know this thing that meant to be an exhaustive list of all the major algorithms that you can use you can definitely look online I included a link to there was a cheat sheet that Microsoft has for choosing machine learning algorithms for your reference or even anything that you want but the statistic kind of gives you a sense of where topic modeling falls in this well and this isn't no because Lda makes use of cluster million variable structures if you're not familiar with that clustering is there to reduce the number of examples and into parser and pretty much it generally depends on some sort of distance measure right so points near each other are normally within the same cluster and points far apart from each other and normally in different clusters and the reason why a lot of times is paired with dimensionality reduction is because in high dimensional space distance measures don't really work very well so you reduce the number of dimensions first so that you can select your distance metric can make more sense and also for the second line users today I'll be using a Justin or a Jenson application of LD a psychic learn does have a package and they have an example of topic extraction using non-negative matrix factorization that's powered by LD a that I encourage you to look at but I will be using Java that's right well okay so next we'll talk about topic table now that we know where it falls in the world of machine learning um topic modeling is miscible Paulo for discovering the abstract topics that occur in a collection of data collection of data it is unsupervised learning approach for finding topics and large portions of text and it's really great for a document clustering which I just talked about a bit earlier information retrieval and for feature selection what it is that is not my name so it's not just kind of finding regular expressions or dictionary of these keyword searches it kind of does a bit more than that which we'll talk about as I dive into it some of them are popular use cases of topic modeling in current day for all you marketing analysts or data scientists out there SEO keywords to combat like keyword hacking based in simple keyword matching a lot of search engines now use some form of topic modeling to look for content and revelant our relevance there's also a pretty popular New York Times article trending and how they use topic modeling techniques such as Lee and Erlich allocation for their recommendation ending engine along with like a collaborative topic modeling that takes signals from other readers into account well so I think it's important to talk about a couple of information retrieval methods and some of their pros and cons and how Lda kind of built to solve all of them or how each of them built on top of each other yeah so tf-idf you'll hear that often and pretty much it looks for how frequently a term appears in the document suitably normalized this is also known as the term frequency which is where the F comes from compared to how frequently the term appears in the corpus overall which is also known as the inverse document that you see or idea essentially if a word appears frequently in a document it's deemed as important and scaled up in score and if a word appears pavilion many of the documents counter words like the or and/or then it this kind of scaled down in space tf-idf uses the bag of words method then it breaks down a document into basic word components and I'll go over that a bit later too from that you can then count frequency to characterize so there's some really cool prose about say about yeah you know it's really great for lexical or word level analysis identifying most descriptive terms in a document but some of the cons are that it doesn't really take into account semantics very well or the meaning of the text and there's very little compression of a corpus or dimensionality reduction and it reveals little little of the inter and intra document statistical structure one step further so after your tip idea latent semantic indexing or LSI it takes tf-idf one step further by capturing linear combinations of tf-idf features and using singular value decomposition to perform dimensionality reduction on tf-idf words or vectors and capture most of the corpus variants some of the pros of lis and semantic indexing is that it achieves significant compression of large program of large collections so yeah so does she seeming to take the question of wide collections thanks to the way that it captures the linear combinations and it could also capture things like synonyms and policy mean and policy names is pretty much like when a word changes depending on the context of the words that it sits with so how like I think I was talking to alexa earlier and we're saying you know bank or whether you could bank on it or a bank account and things like that so yeah but some of the cause were or some of the cons are that singular value decomposition is very computationally expensive and it usually needs to be combined with tf-idf and it can't quite operate on its own and latent semantic lego semantic indexing was kind of flipping about earlier with the SEO hacking and it came as a solution people trying to cheat search engines by cramming meta keyword description tags full of hundreds of keywords these contents will nothing more than random key words and no subject related matter or worthwhile content so you'll find a lot of yeah just Studies on that one so peel the LSI or probabilistic latent semantic indexing allows several topics per document embarrass collisions so that each word will get some topic drawn from the multinomial distribution unique to the document sometimes you also see probabilistic latent semantic indexing or PSN pls I referred to as TLS a sit there yeah so some of the terms of key OSI is that it's a lot more expressive than the regular latent semantic indexing and it's really useful step towards probabilistic modeling of a text so it tries to take what some of the shortcomings of tf-idf and william semantic indexing to the next level but some of the cons is that the general the generative semantics of pls I are not fully consistent which leads to some problems in assigning probability to previously unobserved documents so because it's not generative it's difficult to provide an accurate probability to a document outside of the training set without retraining them off the whole model again and another con is that too many foreigners can sometimes lead to riveting so the number of parameters grows linearly with the corporate size in TLS I and the number of training documents so this can sometimes lead to overfitting well so now for the topic model of the our lane deer-like allocation so I'll be a su Mo's documents to produce the mixture of topics those topics then generate words based on their probability distribution given a data set of document LD a backtracks and tries to figure out what what what topics would create those documents in the first place LD is also a probabilistic model which possesses consistent generative semantics and overcome some of the perceived shortcomings of TLS I derelict distribution is just a lot of questions about dearly but during the distribution is a probability distribution over the space of multinomial distributions pretty much the spice there with prior bills intuition into the LDA model that the document only covers a small set of topics that use a small set of words frequently and this results in better disambiguation of words and more precise assignments of the document topics which is why a lot of people like LDS in general the dirac like I said the direct part in LD a can therefore be interpreted as a regularization method like l1 normalization like an l1 normalization is pretty much taking the absolute deviations or errors and it basically minimizes the sum of the absolute differences between the target value and the to value another Ellen regularization technique that some people might be familiar with is the lasso method as well ah they also overcomes the overfitting problem by treating the topic mixture weights as a key parameter and K we'll talk about like what some of these variables mean in the next slide when I go over some of the plate invitation but K is pretty much talking about or referring to topics but it takes some treats to topic mixture weights as a key parameter of hitting that random variable rather than a large set of individual parameters explicitly linked to the training set l da is a generalization of the pls I and generative in the sense that it can give a probability to a document outside of the training set yeah and unlike other d reelect multinomial custom models l da does not restrict the document to being associated with a single topic which depending on what you're analyzing or how you want to use the data that can be really useful so one of the examples that I'm going to go over a little bit later is whether I brought it before the turn of these insight and so there's a couple different ways that you can think about churn you can take this rich free text data and you can shy and aside one topic to it and you know categorize it that way or if you want to get a general understanding of the entire purpose you can see like for each journal easing what are the top three topics that are that come up in this trend reason and then kind of averaging over the purpose to kind of understand like throughout the whole corpus what are some of our top top topics and top things that people were writing about because sometimes as we often you realize if someone says let's say I'm leaving the product because you know I earlier for the price of X amount per month I really would like this feature and x y&z and so sometimes it can be cost in combination with like a feature across you can categorize into one or the other but sometimes knowing both is is helpful depending on what the product ones so it's really cool that all da has that flexibility well so through the Flint notation I was talking about earlier this a lot of times we use Planet ation to visualize what's actually going on behind the scenes of Lda and how the main variables or the many variables are related so in here the boxes are plates or representations the outer plate represents the document and the inner plate represents the repeated choice of topics within a document or topics and works within a document the Big M represents the number of documents and the N represents the number of words in a document and case refers to topics so this is really important to understand because when I start throwing through like this walkthrough or the pipe that walks you through all da you'll really need to know what these parameters are it is going to tweak it as is every classifier you can always they will always be default settings for like alpha and data but the more you get to kind of understand these parameters and like how they can affect the model the the more you can kind of like to keep your model to to better fit Vidia okay so alpha and beta concentrations alpha is a parameter of indirect fire on the per document topic distributions and beta is the parameter on the topic word distribution alpha beta concentrations are parameters in the model that you can really tweak if you have some domain knowledge in particular especially about the the model and the product or the deer that you're working with so the higher the value of alpha the more topics documents are composed of and the lower the value of alpha the fewer topics on the other hand for beta adjusting it if you're just baited to or the higher the beta the more words in the purpose that the topics are composed of and with lower values of ADA yours so just kind of keeping that in mind and so if you look at this you see here that beta is a popular submission for document M and bar Phi is the word distribution for topic K so if you're looking at this plate location you can kind of see the the other alpha and beta are the parameters the internet beta in bar Phi R vectors storing parameters over or parameters of the Bureau of distributions in LD a words are the only non lien variables that are observable so everything else is hidden and as I've mentioned a sparse day with prior which usually equates to like alpha being less than 1 is put over the topic word distribution which codes the intuition that the probability of topics is focused on a small set of words okay so we're almost through Barry I know it's really deaths but I think it's really important to understand before I jump into the code and kind of show you how this comes to life so the last thing I'll talk about is mala valuation metrics and they're just like a couple of different ways that you can evaluate your model and just make sure that it's measuring what you actually wanted to measure and that it's giving you output that you can make sense of so one way or a popular way to kind of like try and measure the output of your model is using the human in the loop so duo this normally seems like a word intriguing their topic intrusion and that's pretty much when you let's say that priced and feature are the topics of these these are the topics and cost paint money support iOS Android scheduling annual are words that on these topics right with word intrusion you would swap a word in one topic with another and if a human or a person can find the insurer or the intruder word and the topic is done so like for here for instance under price to be cost pay money and then support support here is the attention to network and hopefully in topic is strong enough you can pretty easily tell that and you can do the same concept with topic into gym but sometimes even a loop community possibly because you have to actually have people so these are the ways that you can evaluate your your model so other ways cosine similarity and pretty much this is calculating the similarity between different documents by computing the intra distance and intra distance between vector space and it seems that topics are spread evenly so if the stores so you know that like the topic is good as it has a similar scores if it has a related scores or opposite scores and you can kind of use that to guide how your topics are doing and and how the model is performing a third model evaluation metric is predicted perplexity so pretty much algebraically this is algebraically equivalent to the inverse of the geometric mean per word likelihood and a lower perplexity score indicates very generalization performance so you're trying to optimize your model over this and this is kind of a way that some people have figured that they want to choose their topics or if they can't really decide how many topics they want to split their corpus into this is a good visualization that helps you about all right great so now I'm just going to dive into some Python libraries that I'm going to be using and my walkthrough so yeah just to familiarize you this is these are some of the Python libraries I'm going to be using NLT kay is a natural language toolkit for Python it's newly's full package for any natural language processing as it only will probably tell you what Nina's caucus ball there's also stop words such as a Python package containing stop words and stop words is pretty much remember I was saying those words that really don't add meaning to her sentence like the or your most stuff works of following like that so summers packages or package will kind of like take those words out and indent them which is the topic modeling package that contains our Lda model okay all right so a couple of major things that i'm doing and i walked through that i want to kind of just go over now so i can breathe through it worth showing you but one of the major things there's a bunch of common steps in most natural language processing methods and the first is to predation and organization pretty much segments the document into its atomic elements so say you have i went to the mall it's combate sentence is composed of a bunch of different words and so the organization will split that sentence into a token or each word and so in this case we're interested in tokenizing words and we're going to use NLT K's tokenize resurrection module to match any word characters until it reaches a non word character like this face they're actually baptized tokenization and there's some like workarounds for things you get like interesting things and you'll see in the output like T and you'll get T because when you're tokenizing it all like split like a contraction for instance because a pasta she is an on board characters when you get do n and in T and that's just something that you'll just have to be aware of this kind of so knowing visualize visually and there are ways to work around it but just something we should get we're stuck words like they said so certain types of English speech like conjunctions like for or or the way that remedial is to a topic model so these terms of cups up words and should probably remove from my trip unless we use the stockers package from pi pi in my tutorial and it's relatively conservative it's a relatively conservative list there's a lot of different software packages that you can use cited for you to explore a little um the last major pre-processing that I do is stemming and so I think stemming is pretty important because stunning words is another common natural language processing technique to reduce topically similar words to their route for example stunning summer sun wall be reduces to step and this can be important for some topic models because sometimes they would otherwise view these terms as separate entities and that could possibly reduce their importance in the model so I found that by doing they help to improve the accuracy and relevance of my top topics oh and the last thing I wanna talk about is document term matrix so pretty much all the text documents combined is known as a corpus it's just so from same clip is what is the purpose nice no to run any mathematical model on text on a text corpus it's a good practice to convert it into a matrix representation so LD ale-8 model looks for repeating terms or term patterns in the entire document turn matrix and to do that we need to convert our focus into a document matrix using Jenson so the the corporate module assigns a unique integer ID to each token while also collecting word counts to develop into statistics so like I said the other ones kind of reduced the words too immature consumption was active and accuracy and based on stemming it reduce it down to zero so it can get after it accuracy after you what however or else you can make that word and then it assigned an ID of 657 so that it makes a word to an ID in the document term well so it's my daughter's now I'm going to show you an example of using Lda and how I use that in my role to help analyze to oh um so yes I can actually see this thing [Music] it's funny because they say um they say AI is a AI is easy and a V is hard huh this is China let's see one sec alright so this is pretty much a breakdown of kind of how I put everything into practice this is kind of going over some of the things I've talked about earlier saw kind of greasy wet but he's just importing the packages that were there unions I imported document and essentially wrap my sequel code that expects the documents for our database in Python I need a bit of cleaning and reprocessing just because for some of the show reasons they're like automatically generated within our company and not knowing our product I know that some of them aren't very useful so I don't necessarily need a topic for them if they're automatically generated I'm looking for I'm looking for sure I'm give me one set that's better a lot better cool all right so yeah so some of the cleaning that I did at top was just kind of understanding what's not a benefit generating but I actually care about so and don't don't worry about I guess taking pictures I'll send out a link to one of the organizers with the presentations that you can have this on file and the snippets of the Python code are and the representation is law well so after I cleaned up my data set I transform it to list so that they better speed it in later did the tokenization in stock word since coming that we talked about these Porter's better okay this loop essentially from that traitor to lube to feed it down memory or my data documents through so it just gave the Lucchese tokenized remove the stop words ten words and then add the two these two lists here well and then after that I've constructed the document turn matrix so like we talked about the corporate corporate module I'm assigns like a unique integer ID to each unique token while also collecting word counts involved in statistics like discussed and then it our dictionary has to be converted into a bag of words do those great and so the next step after you kind of do everything that you need to do from a natural language processing perspective now you get the created object for the elia model and train it on the document turn matrix the training requires a few parameters as inputs which are explained below and when using the Jennison module so some of the Vanaras that you must input is 1 the number of top number of topics so just like clusters you have to kind of tell it how many splits you want or like how you have told something I just can't do it automatically so same with Elliott tell it how many topics that you wanted to give it me being a product analyst I have kind of an intuition of the product and so I kind of know that maybe there probably will be 15 major topics that I'm interested in started from there by actually I start from 20 and I kind of like dwindle it down to something until all the topics team media enough and relevant enough to us in our product to be able to understand and another thing that you need is the idea word so ldea model requires the previous dictionary to map IDs to string so that's we've got the document and matrix comes in and pretty much the audit award would be the dictionary that we defined that map's like for instance accuracy is 657 just so that it knows how to map it whatever new document it gets and then the purpose would be the curt the new corpus a document that you want to chain 1 and chain your model 1 and then passes so passes are optional they're not actually required to make this thing run but the number of locks the model will take through the corpus it's pretty much the patent and the greater the number of the passes the more accurate the model will be a lot of passes so can really slow down things on a very large corpus so just be aware of that kind of similar to like some of the like cross-validation and now like depending on the split stuff for everyone so just keep that in mind and there's a lot of other things are included in your grammar that you can't change like I said there's the Hydra parameters of alpha and beta that you can change and if you look in the gemston model or the jensen documentation but if you look in the Jensen documentation it kind of tells you what some of the default settings are and if you want to play around with them like Alf I think in symmetric you can make it on a symmetric and a couple of others speaks to it so you can play around that and see how it improves your model well all right so after all that now we have it is trained and is ready to go so now we can review in the topics the lab so pretty much I wanted to be 15 topics and finished topic I'm on the scene what are the top three words associated with that topic so the when you're seeing here is each parenthesis is a topic with individual topic terms and week so for instance this is topic zero and useful and current are some of the top words that are found in that topic and etc it's just in the models number of topics on and past this is important to really get a good result and excuse me eliminating confusing topics and because like sometimes viewing and in this space can sometimes be a bit dense and it's very hard to visualize luckily Lda has something called paul bivas which is a really cool tool to visualize the fit of the LDA topic model to our original purpose so pretty much if you pass it through but your modelling put the purpose and you put the dictionary and you visualize the data and pop database gives you this really cool visualization of pretty much the top hooks one can say so these are the 15 topics that I defined based on model and essentially the larger groovy bubble is kind of like the larger the over the medium the topic the smaller one or smaller the bubble those kind of like the smaller the topic where you're seeing here to the right at least is some of the most development terms for topics so for instance topic number you'll see terms like service just you think time and this may not necessarily make sense to to you but knowing my product and knowing so the things about it I can tell what topic this is and know how to categorizes it know how to kind of inform my p.m. on on how he should proceed with analyzing attorneys and that kind of fall within this just so that you got is clear as to what we're seeing here the blue bar is the overall term frequency and the red bar is the estimated term frequency within a selected topic you have like a lambda bar where you can like slide to adjust the relevance normally a sweet spot anywhere between 0.5 and 26 and really what the relevance the relevance metric is pretty much saying one shows the most popular purpose white terms and 0 shows the most distinct terms to each topic because remember Lda is looking at topic 2 word and topic topic the documents and so we try to find a happy medium in between and you can adjust accordingly so this is yeah this is kind of like the output of Lda and i mean if you want to see it like the way there was there a couple of different ways that we applied it but essentially one way was to look at the top or the top predictor so for each document and in this case document and there were a couple of different topics that were defined in it and what I did is I took the topic with the greatest probability and just assigned it to that Treves and so what I would tell this journeys in topic 2 so I didn't know why Rizzo so that's probably like accidental build next billion to one of the cases description too expensive which would be talking 5 etc there are other ways that we kind of explored this data so we looked at you know for each turn these in what were the mix of topics are the top three topics and then we looked at the we kind of like averaged over the whole corpus so that we can kind of understand okay we don't need one chair reasoning to necessarily be one topic you want to see in general what are people talking about what are they concerned about and if the turtle reason includes like a couple different topics we want to know about that too so yeah this is this is all the a are there any questions yeah how much do you need [Music] that's a good question so I mean of course like oh yeah so the question was how much data do you need to have the topic Moll and of course like with every every input and output the more the better you need a sufficient amount to be able to for your your mouth to make sense of it and it really depends on how meaty your documents are so for this my vacuum is like doing that maybe some of them are some people really like to write a lot so they do I like paragraphs you know and I don't know if we have a word limit maybe 255 characters standard but yeah but for some some uses an Lda like New York time for instance like their document as a whole you know article and so it'll be a lot meteor and so I think you just have that understand how robust each of your documents are and then how robust the entire purposes and see and just kind of getting a natural intonation of whether or not this will be able for these results that are meaningful to you and if all those photos try it try it and see and see what it produces and if you knowing what you're working with if that doesn't make sense to you then you probably get you that and tweak the parameters or you might need more data where you change the Wikipedia and then you can I've definitely seen some papers and articles on that but I think you just have to make sure that it's super bust the computer article and so like if it's really concentrating on a certain type of jargon and only has like a limited amount of words to trade on this will that might affect like how accurate it could be different purpose but yeah he's really good a lot of times I think there's like the New York Times set which when you're learning how to do this or some of the tutorials may like to take you through that salt line you can look at it yeah instead of treating the document as a bag of words do you think it would make sense to extract the subject of each sentence and feed those into an Lda so so instead I was like looking at you so if you wanted to so essentially you're extracting the subject so I guess I'm going to use because the subject would do the topic right so in this case but what I'm thinking of is if you used some natural language processing I don't get to take each sentence and figure out the subject of the sentence [Music] instead of behind yeah that me I feel like that that would definitely be interesting to explore but I think a lot of times with Lda you're trying to get the topic which is kind of like the overall gist or subject of the sentence and so at what it what what ld8 does is that it looks at like all of these different sentences and it tries to figure out based on the words that this sentence is surprised of what is the subject of the sentence so it's trying to figure out like what is the topic or the overall gist of it not if you're saying only concentrating well I thought oh he was you have a document with multiple sentences news so instead of taking a document with multiple sentences I'm just reading each word as a completely independent thing take each sentence pull out the subject and use the subject of each sentence to figure out what yeah and so that's definitely something that I just can be done it is yeah I mean I've seen like a couple approaches where they don't use bag of words and like kind of treat everything separately and so maybe that definitely can't be done it just wasn't the approach to I accept at this time but that they are interesting avenues we report coming up with topics labels can be kind of tricky and you take like one example from the visualization and just kind of go through your conclusion yeah sure okay so these on this right just from my intuition because like I said I am a product analyst and and the way that we do product analytics at least squares that we embed our product analysts into the product so that they can really be owners of all things data and analytics for the product but it also is a great way for them to get natural intuition about what some of the pros and cons of product Neutron shortcomings so when I look at this so service just thank you use time season feature now please so based on those top words what I get from this is that this topic is talking about some of these support seasonal and seasonality is like one of the reasons why they churn some of the it could be a combination it seems like in this topic there was a lot of combination of people who they said that you know I would really like to use it I just can't use it now because you know for this feature are we need I already do speeches or this particular time or something so yeah so it's a little mix of that and just look from my intuition of knowing the product I know that those words kind of go into it and - I guess to make sure that my intuition is correct I have this output at the bottom that pretty much all the sign like say if I only wanted to look at wine and Charlie's in all the signs of this all exported to UCSB and I'll check this picture that I intuition is right so I'll look at that how we'll get like all the actual chart reasons that would be categorized there and say does this make sense I'm sorry be perfect it's definitely gonna be some that aren't necessarily the best fit but it's useful in giving our product manager this kind of a sense of what's going on yeah my question is actually related to this one the number of topics is you set the parameter in the bottom very much I never cry yeah and just for me knowing it's also for me I start a little bit higher than I like work my way down I just I really like elimination that's like a lot of my techniques especially even with machine learning but yeah knowing the product that knowing I guess I could splice as I say to a lot of different things but how useful would that be to my to my product manager to have like 40 different topics of why people are churning so I kind of knew that anything maybe 20 or below would probably be granular enough to get them a detail of what's going on but not too granular where it's kind of lumping things together sorry 720 the bigger topics were great some of the smaller ones in a room make sense or uh yeah they make sense so that I I work with 20 and kind of like work my way down but there's also a really good paper by state color let me see put that up here yeah callback lie blur division score toss to you if you look up that up it'll be a good way to see how to find the optimal number of topics or your your purpose Nemo so for you with that as a product analyst I'm sure you have a preset classes of types of things either looking for what's the pros and cons of using unsupervised person is supervised by I'm if you kind of know what you're looking for before you yeah so there are another good question but I just for the data that I'm dealing with here I do even though I generally do know that the types of issues I am looking for because of the way that our data is there they're not matched it's all of our data is just like three types three types and even though I've you know let's sing the output for instance I don't have the output legal right now in in our database so first step could be labeling the output using Lda and then second step could be like once we have a match input and output then we can use some type of supervised learning for future cases but that's a good step that's a good question for sure [Music] yeah and so on like I mentioned let me see if I can can pull this up so psychic learning has a really good example let me see Lda using [Music] I think I have a link to it in my presentation in one second [Music] well so yeah so second one has really good example of topic expression with non-negative matrix factorization and like nearly allocation I personally haven't gone through this as of yet too personally Mabel at all like which one was better for my needs based on some of the articles I read I think Jenson was a good first step like getting a basic understanding of like what's going on with my data Jennsen seem to be really good for that trying to get more nuance in the standings a lot of people refer to this as like a next step but yeah I would employ this is oh I have a link to it in my presentation you can take a look see what the students going on the top extraction with non-negative matrix factorization and Lda and really completely compare the outputs from what I've read like it varies depending on what your your purpose what your purpose looks like that you're treating it but also like the output that you desire so unimportant a test about yeah [Music] I have not tried it on different languages in English probably only because the other night answered I know I was Italian and I'm sure we have soft words in Italian so I'm sure that they can you can definitely do that maybe that might be a good question for Antonio when he talks about math and natural language processing if you can tell you about how that works in different languages I have not tried it though most of the charities is that I've been analyzing happen in English though so that's a good question but I do know that there are natural language processing tools for other languages so I wouldn't think that it's impossible I just wouldn't be able to tell the topic so let's see yeah learning kind of topic modeling applications out there a few guys where you don't have to go out and take all these pieces of the puzzles and put them together more about that someone else have to kind of put that all together and the documents and you know I wish all our I wish if you could make a perfect package like that by having everyone in here they love you forever but what I do understand you saying like psychic learn does really good job is like trying to collate like a lot of things together right now Jenn's to them they have a lot of really cool tools I don't think that you have one place where you can do it all I haven't seen a package that pretty much has it all if you do you scikit-learn you might be able to just get away with using you know like importing a couple of modules and and using those but yeah so those are the three major ones that I could really recommend alright thank you so much for you [Music]