SF Text: Oleg Rogynskyy, Q&A with Alexy Khrabrov @Nitro
Recording: SF Text: Oleg Rogynskyy, Q&A with Alexy Khrabrov @Nitro
uh hello everybody i'm alexi krubrov the chief scientist at nitra and we have a nitro tech talk series and as a part of today's tech talk we had a guest allegra ginski who is the ceo of simantria alexa lytics company uh and now it works in the nlp field uh that actually connects to the sf text meet up with sharon and text by the by conference who met in the context of text by the bay and uh oleg uh run successful business in the space and it's always great to see where nlp is applied what are the market forces driving this so uh i welcome oleg it's great to have here can you tell us a little bit about semantria and like solitics and what you've done there oh i'll take a step back my my experience with nlp goes a little bit further than that i was one of the early employees at a startup uh called einstein technologies which was out of montreal in 2007 which was sold successfully to open tax corporation in 2010 after that i joined electrolytics which is probably the last independent last standing independent um kind of commercial nlp uh engine out there and in 2011 i spun semantria out of exotics so i used there nlp engine which is kind of similar to stanford lp in many ways but it's not open source and there's been a lot of kind of private heavy duty r d put into it uh over the years i spent semantic out of it to build a cloud solution out of it and so that has been really successful we cement right now is probably the largest cloud nlp solution by the amount of data processed daily so that was going well and in 2014 i sold semantria back to electrolytics wow that's uh that's a great trajectory yeah uh so what uh motivate you to get into this space in the first place why nlp why text because there's a lot of it okay there's too much text right and so um the initial step into nlp was with publishers we realized that i i was working with bunch of publishers in fact my first job out of college i was building a website for a publisher in boston and i realized there is so much content it's just impossible to manage it all so instant technologies was around managing content for publishers then uh i came to like solidux on a premise that just classification is not enough uh there has to be other dimensions to text and so let's score tech is sentiment analysis where with sentiment we start adding colors to a black and white picture of text that was going well and now we are getting into four ways of intent analysis and classification kind of more automated ways to classify and cluster content more advanced entity extraction across languages across domains as well as contextual understanding of uh there's a difference between a patient in an er room and a patient person who is really nice to you that's right that's right so so those are the challenges that we run across every day and the funniest thing if you go to a business user the sentiment on the street is sentiment analysis doesn't work and it doesn't work because there is no contextual understanding so you have to actually bring it back to context in order to uh apply it to real life business problems yes you know i must say everything you said is just sounds like music to my ears because because this is exactly the kind of problems that we need to solve at nature of a large scale document understanding right and and it's very interesting how you talk about intent because text by the bay it's about all kind of text and you start with the publishers it's a part of the text pipeline you cannot just do an lp because if you do it at scale you put put it in the web scale harness you need to know how you process all this text yeah right and and intent is like kind of ai because it's behind this little string of text we eventually want to discern the entire this is what google does you know you search for or be os and gives you a ticket from auckland to boston right so like what's the intent so so it's really fascinating to hear that you're thinking about this in the same terms right and the same series of tasks actually as kind of constitutes this evolution of this pipeline yeah uh so maybe you can since you're uh immersion in the business world and you deal with the uh uh customers uh can you talk a little bit about what do you see in the market who are the users of these tools so um the biggest uh market share the biggest interest in the market we see right now is coming from all kinds of listening tools so i see the whole text analytics world kind of evolving from learning about what's in the text to understanding what's in the text to acting on upon what was learned and understood from the text so right now we're kind of in v1 of of the whole nlp world where everybody tries to listen and we get into a very mature stage of listening uh applications so we see large social media monitoring applications such as sprinkler spread fast hootsuite etc while listening to billings they listen to full twitter and facebook firehose in real time and pulling out insights yes then we have news monitoring guys so the ones who are looking to at press releases news etc guys like trent kite uh actually in this building upstairs there is melt water and so on and so on so those are getting big and uh we are seeing more and more people looking into what's inside the enterprise so people analyzing emails documents like what you guys are doing uh for different purposes from litigation which is e-discovery space to document management where you guys play to kind of uh employee satisfaction and so on and so on so basically there is you just mentioned like there is a significant number of different players who need to understand tax uh if you look at uh interaction with these companies which forces do you see inside the companies what are the who the people are talking with it's product managers manager so usually it's product people or top management that realizes there is too much information that information is not actionable so they come to us and and they figure out what more the which other dimensions to the attacks they're not seeing that they can report on after its process with nlp tools okay so uh but in order for some projects of this scale yeah to go they probably should be buying from the senior management of these companies and i'm so i'm wondering uh uh who are the senior manager types do you see who usually drives this what like can you kind of uh see is there like an archetype of an executive who actually cares about understanding nlp and and how much do they know about technology how your interaction it's data scientists it's head of data sciences it's head of product and it's very often actually in more technical companies the ceo himself uh so those are kind of our ideal titles that we're talking to but usually it's the mid-range product managers are the ones who have to actually deal with uh deal with the mess okay how much do you see they already know how much do you need to educate them about what's possible how do you kind of match their expectations towards what your product can do yeah uh usually like in the bay area a lot of people know what they're talking about um east coast is catching up when you go to europe apart and canada it's pretty much behind to be honest we're seeing a lot of activity in brazil okay which is surprising they talk a lot yeah exactly besides that in the valley i'm seeing a very strong bias towards machine learning more than nlp and it's very important to kind of distinguish both they're different yes uh and so there is kind of a religious bias here that machine learning can fix everything which i can i can throw hundreds of use cases where machine learning is going to be so much more work than kind of pure play classic nlp yeah this is actually it's interesting you make this distinction because uh i mean uh i went to grad school at panorama one of the strongest groupses and i've seen how nlp drives ml and usually top nlp researchers are excellent machine learning researchers but there is a lot of craft and there is a lot of specifics right and so i think it's it's actually great to have uh understanding of machine learning but it's not in itself sufficient right and that's not the only thing yeah the most powerful combination is machine learning combined with an lp so you can do a baseline on a rules-based nlp rules and dictionaries and then what you can't get easily with nlp you just teach the machine yes so uh you mentioned uh several geographies which kind of connects us to this question of multilingual analytics and obviously if you go to europe or asia you need to have expertise are you operating globally and what do you do in this markets how do you go about getting top nlp expertise in these languages so uh that's kind of a bread and butter it's not hard to get english language nlp you can go open nlp stand for nlp link pipe you name it but the moment and a lot of our customers went open source route beforehand and then the moment you need to scale globally that's when you run into problems because there is very few open source language tools that support three languages and when you need 15 that's when it's going to get very expensive so our core kind of business preposition is we do the same quality and i'll be in 15 languages right now and there is more coming so we'll have 20 by the end of the year and the way we approach it is not just to give us a training set and we built some kind of random forest model and here you go we actually hire linguists in every language so native linguists in chinese in in korean in japanese in russian etc we break down uh break it down to part of speech tagging subjects of object tagging we uh build out linguistic rules that are unique to each language and we build our dictionaries so kind of the core uh the the basement of the nlp part is done um by professionals trained to do this and then we obviously hire a whole bunch of contractors in each aspect of geography to enrich dictionaries to kind of critical mass size yes i'm not just actually dictionaries approach works well in streaming context so we had a recent talk at lithium which was also upstairs and they processed twitter fire halls essentially and uh used to be our customer oh cool yeah yeah and and so that really works um for instance much faster than uh stanford and er right because you need to do it in a fraction of a second yeah so a lot of times investment in this carefully curated dictionaries pays off for for streaming analytics exactly uh so is your platform closed source uh how how does the model work and also you mentioned that it's a cloud solution how do customers push data into it and does it it's a simple rest api we've seen people integrate in 10 minutes so you you batch in your json with content you but we give you back so you go and pick up the same json with content and enrich data about 100 different fields we output for every tweet or document is sent to us it is cloud source closed source however we do publish a lot of our research um in the open source community uh we share a lot of kind of foreign language research where we find things for example it's really hard to do sentiment in japanese it's inferred yes we found a way interesting uh same thing in chinese people say oh chinese is better than our english and not to say that english sucks i hope not but again same thing we published a whole bunch of information about how to break kanji into into structural pieces and uh and work with it using a regular part of speech taggers and so on and so on the whole stack on top of it oh this is great i mean i understand like that's a valid business model not everything can be on personal especially if you involve a lot of uh uh contractors who create those dictionaries it's your initial property however a lot of people have their own specialized needs and they have developers so is there a way for people to kind of uh help in creating this dictionary is there a model how do you see this kind of setups uh where you know you have developers you have some folks who who would want to extend your platform yeah so there is three ways of doing it i mean obviously there is the old school professional services which is there uh second one is a lot of our customers come to us and say hey look i have a very specific data domain where nothing else works can you help us like drug detection or i don't know we work with the detroit crime commission they their domain was all the drug dealing slang that kind of stuff so we actually take the data set and we uh we we put our linguists on it we figure out what are the kind of key cornerstone pieces in the data set and the third model which we have with several companies abroad in particular in singapore and denmark etc is people who want to build languages that we don't support yet so our partners actually our partner in singapore went out and built using our sdk and our guidelines our support for malay indonesia and bahasa and singlish so english is chinese mix of english which is used in singapore uh and so we have a partner network that extends our use our language capabilities interesting interesting uh that's yeah i think that's that's a good model uh it's a certain sounds like a interesting way to uh to augment your global reach yeah um and uh so i heard from some people you know sarcasm is a very hard problem right and and uh i have to instant opinion that uh it's very hard for machines such as sarcasm and it's easy for humans because we actually hear it we hear the sarcastic tone of voice so some people think that you know it's not enough like to see the smiles and things like that or sometimes they're not there what do you think about that idea that you need to in order to understand sarcasm you need to recreate the phonetic the sound of the phrase and you have to bring in kind of speech and voice to fully understand or do is he's not or are you able to do like a lot of it with just string representation of the text well if if sarcasm was done via voice then there would be no sarcasm in any of the books so i don't agree with that sarcasm can be written we've looked at sarcasm a lot it is a difficult problem we've built some technology that haven't we haven't prioritized yet that can tell us there is a high likelihood of sarcasm here but we it's not strong enough it's not actionable yet i don't think anybody on the planet has cracked sarcasm yet yes and the reason for that is that you need to have the whole baggage of knowledge unlike computer you go through beginner school medical middle school high school university you learn the data you learn what's true because sarcasm is is what's the definition it's stating and not true fact as a true fact so the machine needs to know what's true and what's not true to invert it yes so the moment machines know as much as humans then we can talk about creating or detecting sarcasm programmatically so you're talking about google brain talking about some people like watson for example yes and here's another question how do uh kind of uh smaller companies compete with players like watson or or google because google is basically basically building google brain which is the semantic database yeah you really need semantic understanding of a lot of text and but you you probably need something like google to fully under especially you need a world model right not really we we do just fine we for example uh we took a whole bunch of hardware from amazon and we distilled all of wikipedia into a large semantic net okay uh it's a little bit different from what google has because our stuff is curated it's a smaller data set but we know it's curated another example is a company that was bought yesterday by ibm uh called alchemy api one of our i guess now x competitors uh that did similar stuff with a whole bunch of structured and structured data sets uh yeah so we take the same route thankfully with the cloud computing uh we our costs are manageable to do that on on the same level as google and ibm are doing it okay that's great to know so it's uh obviously leveraging cloud computing and proper technologies to build world models uh i think that's that's great that you know companies uh mid and you know small and mid-sized can do that as well that's that's fantastic to hear since we have a lot of startups in our in our community and i'll kind of uh close with uh a few questions about the upcoming texts by the way conference which we're just like connected about so so we're building this new community right where we'll have startups working in the space we'll have some folks from larger companies like linkedin uh airbnb who use this for analyzing their customer sentiment right we'll have non-profits such as wikipedia who actually you know they want to provide a service for everybody and they're looking what what folks are looking for what kind of questions they're asking and we'll have some academic researchers such as stanford phd is working and nlp and advancing the frontier so uh the original premise was that we'll put together folks around open source right we have academics who produce this open source and we have uh uh practitioners who deploy this open source and on their side they feed data and provide the feedback loop to the uh authors right so that's one of the uh use cases which you know we can observe at nitro and other startups i wonder what would you like to see in this community what where can you contribute what do you want to learn and what are your ideas around this you know community going forward uh i would like to see as i'm biased i'm come from business world yes i would like to see more business crowd in it why because i mean for engineers nlp is sexy and it's cool and it's awesome it's fun but when you come to to to a business meeting when you start talking about nlp you see a whole bunch of blank faces so your average business user who is writing checks to fund nlp research is not as advanced in understanding how nlp works so uh kind of bay area community uh should include business people who will be writing checks for the community to strive okay and this is a great uh i think addition to to the set of things we're gonna discuss and you know we're happy to invite you to our conference and be the voice of the business community we'll have some other uh ceos and executives already and hopefully we'll have a panel on the uh business applications and uh real world nlp and hopefully we'll hear some of your ideas there thank you looking forward to it likewise thanks you