Text By the Bay 2015: Kang Sun, Teaching Machines to Read for Fun and Profit
Recording: Text By the Bay 2015: Kang Sun, Teaching Machines to Read for Fun and Profit
my name is Kang Sun I'm with Bloomberg R&D machine learning and today I'd like to go over some of the things that we work on some of the difficulties and challenges that we face and um hopefully uh wet some of your appetites um on on the types of problems that we're we're interested in and the title of my talk is teaching machine to read for Fun and Profit so I'll just go over briefly um and I'll I'll save 10 minutes at the end for questions if you guys have any I'm sure I'm leaving a lot of stuff out um I'll go over the motivation of of why we care about this go over some the problem statements that we have some applications some future directions and then I'll leave time for questions so here you'll see um is the price chart of a very famous top us Investment Bank in 2010 which had a significant price move um during the day so if we zoom in on the beginning of where the price start to drop we see that there was this secc announcement made so usually these these government filing they they have metadata taged to them right and one of the most important things would be the companies of at question right in the document this this SEC announcement was actually mislabeled it was mistagged um the identity of this top us Investment Banking well Bank was inadvertently left out someone in the Bloomberg Newsroom actually read the read the document so there's a period of time when it's published everyone rushes to read it um and so there's a period of time where nobody knows what's in it someone read it they published a headline about it gets picked up by the New York Times minutes later and then you can see what happened okay so shares fell about 10% in 3 minutes this is not uncommon actually we have a lot of um uh instances of stories affecting the price of a security that the story talks about so assuming that you can teach a machine how to detect these sort of stories right um obviously there are a lot of things you can do with this um for with trading risk analytics uh research so forth so trading it's obvious right whoever knows this ahead of time the soonest has the advantage for research a very common question ask in the financial world is the price moved what happened what caused this price to move right so one thing the researchers want to be able to do for building their own internal models is to be able to trace back in time to the event that caused this to happen um so let me go back here so the SEC announcement was is anyone familiar with this case no great um the SEC announcement was that the government was going to sue this bank that's that's what basically happened so if you can teach a machine to read these filings automatically you would have a significant lead for that day I mean could you imagine if you can actually do this very well that's a 12% there's so many many things you can do to to trade on this information any questions about this how was between 12 um this is actually the point in time the announcement was published so it's like 1035 and this is 1040 right here this is 10:55 so actually that was actually I think this is the amount of time it took for someone a human to actually read it okay so this is the advantage you would have if you had a machine doing this okay so the problem statement okay so some some Financial professionals they they want some sort of Black Box some magic Black Box they'll give them three numbers right the sentiment um of of a company in a story the impact of this this news story and the novelty of this news story right um You don't want outof dat information affecting your models so when I say some Financial professionals actually a lot of our clients are not human a lot of our clients are blackbox trading machines okay not high frequency but more event trading uh machines okay so for sentiment we actually have a slightly different formulation I'll go over that later on um is the information and document likely to be interpreted as positive negative or neutral by a long position of and I'll go into that later impact How likely is a story um to impact the price of the of the security that the story talks about novelty is the information new or not so sentiment is actually probably our most uh mature product our area of work um was one of the first things that we we started working on and not only that but we want this to be competitive with humans okay um and this is obviously a very hard problem and people are very interested in being able to do this well for obvious reasons so okay sentiment is a story positive negative neutral for a long position investor so we had to formulate this this is supervised learning we had to formulate this um for for the annotators for the taggers for those of you who saw the first talk in this track it's a difficult problem and so and again the sentiment formulation is a little different from the typical sentiment uh formulations where is this in general positive or negative here we asked the question if you were long basically if you owned the security and you wanted the price to go up is this piece is this piece of is this story positive or negative or neutral to you so it's a fairly well it it seems well defined but it's actually kind of um there areas of ambiguity as well so some of the things that we have to address is how do we combine multiple classifiers together um how do we force this into a probabilistic uh um uh number again the a lot of our clients prefer a single number for sentiment we have NLP issues here named idty recognition text summarization so if you have a a large Story how do you know what parts of the text refer to a to a specific company you're going to have um multiple companies mentioned in the story and we've progressed this to working on headlines and tweets so headlines are basically tweets all right the significance of headlines is this though in terms of um the freshness of information the headlines are the best things you can get by the time someone has has time to write this the body of the story that's we're talking about minutes right so you may be a little late um tweets obviously they're they're that's a NLP nightmare you can't use usual dep parsers um and also companies are mentioned there's a convention on in tweets now you have dollar sign and ticker right that itself is wrought with ambiguities for example um dollar sign f is Ford dollar sign f is also Fiat so you have to disambiguate that as well um people care about Japanese language classifications and then non-equity instruments are very very hard okay so people care about the sentiment especially recently of oil what does that mean is it the oil Market is it the price of oil um is it the oil industry the Drillers the fracking industry what does it mean when you talk about the price of oil what does that mean is it WTI is it Brent it's just a very uh fragmented Market people care about sentiment on countries Greece right um people care about sentiment on macro uh uh indicators as well unemployment are you long on unemployment or or short on it so unfortunately in the past and well unfortunately even today a lot of organizations are inclined to solve the sort of problem with a rulebase turistic approach all right so again we have a lot of metadata on our stories with topics but these topics were meant for humans not for machines um you have lexicons right you have these uh uh lexicons of affect on words but we're talking about a very specific domain area right so what does cut mean what what is the effect or the polarity of cut it it doesn't really apply writing rules is very difficult um deriving the rules even harder treating this as a pure regression is difficult so we do what many people do we use svms or other machine learning techniques as well I'm sure most people are familiar with this sort of thing um this is actually a ranked view of a the most positive sentiment and most negative for a day and then we have some representative headline for for uh the company I was curious how uh this would have worked when is anyone familiar with uh Switzerland having to let their currency float right so all their Banks just tanked right the firms were crashing and I was curious how that kind of ambiguity of currency would affect sentiment and the the day the announcement was made the most negative sentiment had three Swiss banks on there so I was pleasantly surprised any any questions on the sentiment part so the s model that you said for using we have multiple s models sentiment classification how do you cretive training progress okay good question um this is actually so supervised learning we have a team of of people that tag these stories so we have to feed it through our NLP pipeline um identify the regions of the text and then present that to an annotator G annotator and they're just tagging these all the time for for Tweets we're actually using crowdsourcing and and unlike the difficulties of um the first talk here it's a very very concrete question if you own this stock how does this make you feel positive or negative neutr very very very very concrete a we actually do a lot of test annotations ourselves um some of these are very ambiguous obviously some of these are very emotional right let me give you an examp an example when a company announces layoffs in the short term it actually has a positive price effect right but emotionally it's telling you something's wrong with the company right for the long run so if you see a lot of layoffs for the company iron it will have ironically a uh negative sentiment but in the short run it's it's it's inversely correlated with price movement so sentiment for predicting prices is actually different problem from what I'm going to talk about next so Market impact is you have this new story will it affect the price of the security okay so this is again supervis learning um couple jointly with financial time stories analysis so the problem with this is how do you get training data for for these examples right so you need to somehow do enough time series analysis to know what points to look and then to look at the candidate news stories um this is an example of a difficult time series to work with because during trading it's actually kind of sparse you see very few trades it's not very well behaved um keep in mind rathon okay I I'll use that as a for a point later so these are just the um these are the candidate stories for the positive label for for impacting the price this is a more well- behaved security so the prices are are denser it's more liquid you see far more data over time um keep in mind FDA okay and also what I said earlier about uh headlines being basically the nuts and bolts of this sort of thing this is it this is all the data we have aside from metadata so the metadata that we have is both human annotated with um topics and then we have a uh machine annotation process as well where topics are applied by um a mixture of rule a rule based system and then here's another one um again this is a very liquid security this is uh there's some ambiguities here right there's a 2010 oil spill instead of today um let's see already agreed so this may indicate something that's in the past okay so this is a tough problem actually I'm not convinced this is story driven you can argue either way um because this is actually the underlying index which dropped as well so it could have been a broad Market move that's another part of ambiguity here any questions about the broad Market move so that's this is a British equity and um the footsie I think it's the broad British index which is the orange dropped as well so is this because the whole Market dropped or is it because of that story that's that's that's an ambiguity here so okay so going back to um finding sh my pointer finding these points in the time series how do you know what the points of interest are right how do you detect these these time series anomalies for things that are really might be poorly behaved like this um you can't apply well some of the earlier techniques um were to just do a lot of regressions and and try to fit these pie wise segments and then and that's your uh model of the time series that's really noisy doesn't work very well is not time invariant um and it's hard to adjust for diverse behaviors so me as a text person with a vision background I decided to borrow some ideas from Vision which is scale space analysis and you can use this sort of analysis to identify the major movements for for for further analysis so we can zoom in and localize the actual point in the in the time series this is very very stable over time it's um independent of what's before and after given enough a margin um and so sort of the precursor to what we're actually using today it's not the exact same thing um but it's along similar lines all right so for social media um this is also important is anyone familiar with this tweet no okay so this is a fraudulent tweet um I can't remember who they they hacked the associated press's Twitter account and posted this story so for us this would have we we did not detect this okay because there's no company mentioned and we're not working on um we we're not done yet with with sentiment on countries and or Market impact on countries yet so you can imagine this being fairly impactful on the markets right I mean it's it's someone's attacking the country 13 $6 billion were wiped out in minutes so I wish we were able to detect this at the time it would have benefited our clients a lot um even though it's it's wrong it still has an impact on the markets right several minutes later the AP B posted another tweet that said oh it was fraudulent we've been hacked and then it's things slowly came back but there was enough delay there that had machines been able to detect this both the I was ask you think immed of the reaction the ALG this isn't even that immediate it's on the scale of minutes yeah it's not that immediate okay um for um for a lot of other cases like this we're on the scale of seconds so minutes is everyone missed that and they panicked and then they and then everyone went the other way after they they saw the you know all clear sign so again this is a very difficult problem it's it's joint learning between text in in the time series so we generate the candidate stories okay that in itself is not enough we actually um do this in a supervised way these are also annotated by humans you could have very this this problem is difficult diffult enough as it is in their literature half half the author thinks it's impossible the other half is they think it's very possible we want to eliminate as much noise as possible um so after we find the candidate stories we go through a round of human annotation and they annotate this as yes or no they think what and that that requires a little bit of domain knowledge right to be able to say yes the story was a relevant to the price movement or relevant to the company and relevant to the price movement so that's that's a challenge for for annotation this probably W won't work very well for crowdsourcing because it will require a lot of domain knowledge training um I wish it would work well it would make my life a lot easier but um we're looking at probably more crowdsourcing venues with more um certain skill sets right like NBA students or so forth okay so for social media we also do work on sentiment um and we also track the the anomalies of the number of tweets over time right so we call the Social velocity and users can set up alerts on um the abnormal tweet be tweet Behavior or tweet act activity for for a company again this is a very difficult problem you need to properly tag the tweets as talking about a certain company right in general our clients care a lot about Precision so we sometimes will sacrifice recall um it depends on the application so it's it's it's you want to adjust Precision recall for your audience right in the application um we have other applications where things like search you know you want recall to be a little bit higher okay um so question and answering we have the beginning prototypes of a of a question answering system we have a lot of data structured data that's really hard to access the the application teams heroically try to build uxs uis that are that makes this whole process easier but still it's it's complex data it's very hard to to grab and not only that but it's Federated data it's all over the place different domains so one of the things that we've been working on is converting natural language queries into structured queries and then after that we're going to work focus on a true question answering system but but the first step is how do we just access the data that we have I mean we have a a wealth of data so simple little thing here all stocks that have their price go by more than 10% over the last 6 months for someone that doesn't use the application it's it's kind of hard to do so we convert that into this query and that's it this would probably take me five minutes to figure out how to to to express in this in this application um this is an equity screener so we do the same thing for fixed income which is actually has which has other subtleties uh credits score um so forth um we're working on several other domains right now that can't go into but it's it's something that's going to be really popular we joke that eventually the Bloomberg terminal is just going to be this one text box you put in some natural language query and it gives you whatever you want mainly because we have well we had over 30,000 functions so it's really hard to find stuff admittedly um and we are as a company always working on simplifying simplifying the process of discovery of finding the data that you want of um customizing searches for you so unlike the large um search companies we only have about 350,000 users or 400,000 users we can build custom rankers custom recommendation models for each user that's actually the focus we're taking we can personalize the entire experience for for a user how many how many users about 350,000 330,000 paying clients um those are the terminal clients and then we have the the non-terminal clients as well so you can go to I think bloom.com about or something like that it has all the stats and then you add in all the the employees and so forth and okay so feature directions um this is just sort of a a landscape of what we've been working on um so when the team started like 0809 these are some of the projects we started on the Senate analysis started first we have a lot of um we have a lot of using off the-shelf NLP models is hard for us because we have very specific Financial language so we train a lot of our um our NLP models internally based on our own or own corpor um let's see moving on you can see more projects are coming online and more this is 2012 I don't know if I have an updated one 2014 and we have a lot more non-text stuff going on um this year as well let's see so future directions there is a lot of room for us to explore we have again an immense amount of data we haven't even looked at whether it's structured or unstructured um cross lingual models are things that we care about as well for example we have a lot of training data for English building these sentiment models for non-english is very difficult because we just don't have the data for it the training data for it so um there's a lot of research going on on uh inducing models based from from another language uh the the most I guess popular topics of the day word embeddings sentence embeddings paragraph embeddings deep neuron networks um are things that we're looking at um so going back to what I mentioned earlier on F FDA and Ron there's a lot of background knowledge about the the name entities we're not taking advantage of as well um and from the notice that I said earlier on some Financial professionals want these blackbox models the rest they want very interpretable models they want things that they can go to their bosses or Congress or whoever and say look this is what our models told us um this is why we made these decisions okay so this is a actual common theme for some of our clients and the other clients they don't really care about interpretability be happy to take any questions for the input to support Vector regines how important is parsing to that part you're using words you're actually using negation or qualification um we process over a million stories a day our whole NLP pipeline has to finish I don't want to get specific to tight numbers under 60 milliseconds for for the sort of thing so we do shallow pars right now um the from the Keynotes I think the keynote speaker mentioned parsing hasn't been funded for a while we're actually funding parsing work right now for for fast parsing because we need the real time aspect and it does help with the accuracy enough to matter the part of speech tags definitely are important the part I'm curious about um for other things that we've played around with yes but it's just too slow right now okay so some of these can be 30 page research documents and we don't want to clog up our our pipeline because again half of our I don't want to say half a large proportion of our clients are these real-time subscribers any other questions is this uh question answering system that uh you talked about is it in the hands of customers right now the equity search is um some of the other types other domains are in beta right now and the feedback has been really good um and then Mike Bloomberg loves it um what's the so you're you're pretty much like taking like a national language query and tur into some sort of database quer right so there's a lot of semantic parsing going on there's a lot of parsing parsing going on so and with that one you're okay with the latency being higher right like no it's actually has to be really fast as well because we also do autocomplete and um if you can if you think about autocomplete going to a terminal in this client server model it has to be really fast or people on the other side of the planet are going to see these high latencies so can you get away with shallow parsing with we have a um custom built semantic parser that's really really fast the guy is fanatical about speed it's kind of like customized for the kind of queries that you it was built from the ground up for this specific purpose um how how do you train these sentiment like Raiders or do you just find people who know the sentiment and no we they actually are um um on staff we have a team that deals with data and we train them there's a um there managers there who train them continuously we um we do all the tests inter inter annotator agreement the quality of a specific annotator um we periodically tag them ourselves and see how how well they agree from from our perspectives we have people on the product side the product managers they tag them as as well um so something to keep in mind we for us Glamour and deceit is not possible because our clients are constantly back testing all the time so if they see any change in quality they will let us know this so for the price impact prediction are you only looking at the impact on the and or for example are you looking at the impact on B um so that's something that we're work there's we're working on we have again we have this wealth of data we have for example the entire supply chain relationships between say all of Boeing's customers and all of Boeing suppliers so we're working on models that will propagate this for example um Apple sales impacting Fox con right or people looking at foxc con's output it's going to impact Apple so these are highly related um this is this an ongoing it's ongoing research right now any other questions where do you get the supply chain information or Bloomberg we have it all I mean we have premium data we have public data we have government filings we have government data we have all of it there's an entire Army of people that works on automating the ingestion of this manually ingesting this um scraping the web premium uh uh content producers we have our own Newsroom um we have our own divisions that work on only government data political contributions uh we have our own legal subunit right so again we are we have so much data that we haven't touched yet it it that's the exciting part this is only on the language we haven't looked at the The Joint modeling of everything else yeah you mention you're using U you show that tweet that how much impact it has so the impact of this social media and these messages looks like these days are it's huge it's huge and maybe more important than conventional so I wouldn't I wouldn't say that yes yes that was billion dollar I I think I think if Bloom this is pure speculation this is my personal opinion not the company's opinion I think if Bloomberg News wrote that the drop would have been bigger okay that's my opinion never came I'm sorry but that message just publish on the no that's just through Twitter no we didn't we didn't publish that we didn't do anything with that right I wish we we did capture that but again we're not looking at countries my question is that okay given that social media has some importance in terms of these signals how uh and right now probably you have some relationship with twetter to get these tweets how how are you going to access these data in sort of real time I can't go into that okay um I'm not sure if I can right now um but we care a lot about social media our clients care a lot about social media um a a good number of my teammates are working on social media right now so the dynamic is this even with journalists things that don't match don't don't attain the quality of Journalism like rumor hear say unsub unsubstantiated things statements it's okay to put on social media because there are no standards there right and then when things become more substantiated and fact checked and so on it goes through the regular journalistic channels so that's why a lot of our clients care about this because I mean that's the rumor mill right um there's a lot of work um you can go I guess for the NLP Community that's been done by the NLP community on tweets right um things like using geotags multiple Witnesses of the same events um you can a make sure that they're talking about the same event make sure that the the geotax Mak sense and so forth and you can say okay these are corroborated uh eyewitness tweets this is a substantial event that happened so that's actually pretty important work as well it's all about events so um kind of along the same lines like except instead of the supply chain um like Sarah's asking if if raon takes a price hit what happens to Boeing's um price we we're um do you have relationship there there are traditional um Financial Quant Quant models for that I mean these correlation models these these these igen baskets of of um Securities um we're we're looking into that from from the nonquant perspective as well so the Quant models have these nice you start with some some very elegant model and then you try to contort it into the data for for Mach machine learning folks we only care about the empirical aspects right so working that working on that from the other direction yeah so I was specifically more asking like how do you identify who Ron competitors are oron like we have that as well um it's you run peers in the terminal tells you competitors and right so that's not some of it's done by human um humans um some of it's done by so the humans would be like business analysts or analysts in general so there there are people whose jobs are just to do that you can get this from um in uh research documents from Banks as well who are telling their clients about the the the competitors of a company or the landscape of a market um again we have a wealth of data we can get there's just so much we can we can do so it's mostly mining or collection that's the hard part right now right we just have our our pipeline as you saw um we doubled in size last year for our our team so we still don't have enough people to to to tackle these problems right now which me do you take to catch the middle to knowa or the market like don't sell too soon or too late like in the case like you so so we we um as a policy we don't give direction we don't say which direction is going to go we're just we just indicate something's going to we predict the probability something happening we're also um because we're providing this data we're also um we we offer trading Services as well it would be a conflict of interest for us to say Direction a lot of our clients want us to give direction but we want to resist that as much as possible because we would have a lot of conflicts of interest internally you mentioned speed was super important for you guys does that change what approaches you use or like what you let your it does I mean we can't use the Stanford parser right um we we care about fast parsing um there's some things that we can do in slightly slower ways that we we don't promise in real time but we want to eventually get everything real time um this is the real time aspect makes it really hard like we're not going to use map reduce for this right by then it's the Market's already returned back to to normal um it does restrict a lot of technical uh it gives us a lot of technical restrictions so well I see the main problems that you fa is that you have lots of data but you have a lot of trouble to get it into a system quick enough to to do anything useful with it do you have any say in how you structure the whole process with getting data in the correct format so it's really fast and accessible for a methods in the way that you need it it's it's very difficult because we get the data from a very diverse set of sources for example a lot of the government filings are not um structured they're just PDFs that that makes it very difficult as well there is no amazingly the government doesn't require you must follow this exact form it just have to have these these sections right um makes things like table extraction really really hard people try to stylize their documents for say annual reports or you know uh the doc the research documents that makes things very difficult um so no in Ideal World yes but no even our own internal um news people we make requests to them and they they politely say no okay all right thank you very much I'll be around if you have any more questions