Devreal

SF Text: Nemanja Spasojevic, Q&A with Alexy Khrabrov @Lithium

SF Text: Nemanja Spasojevic, Q&A with Alexy Khrabrov @Lithium

Recording: SF Text: Nemanja Spasojevic, Q&A with Alexy Khrabrov @Lithium

Hello everybody, I'm Alexey Krabarov, the organizer of SFText. It's a new meetup about text mining, NLP, AI, search, and essentially human intent behind the text. And this is our second meetup on the topic of topics and how to classify content by topic. I have here with me Nemanja Spasojevic, director of software engineering at Lithium. And actually we were formerly colleagues at Cloud where we worked on this similar set of technologies and Cloud is known for the standard of influence and classifying topics of influence on the web. So this is a very interesting area and a lot of you guys probably are familiar with the technology. So we are very happy to be here at Lithium. Thanks for having us

It's great to have you with us. I want to ask you what brought you to this domain? I know you worked at Google before, at Google Books. Can you talk a little bit about what's interesting about this domain space? Yeah. So like I worked at the Google Books for six years and then after that I was looking for something smaller. But still I wanted to be able to have like similar impact like you have at Google where you're working on the like a big scale data. And then like looking at many startups like not too many, many operate on the big data but not like a really, really big data. So Cloud was one of the only startups that actually operated on a kind of Twitter scale of the data and still was like a relatively small like a tense like of the engineers. So that kind of interests me

And then the topical problems are very interesting because like you know like you can start with the very simple things like a you know build like a simple baseline and then you can start to improve like go as complex as you want. And always there is a challenge between academia and industry where in academia you're trying to solve it for very high precision. You're trying to solve it you know to get into the core of the language itself. While in the industry you're basically constrained with the limited resources. And then like it's a kind of very, very interesting problem. Right? Like there are two different sides of the problem. I remember vividly you know Cloud was fairly constrained. So I wonder at Google you have the support of all the engineers and all the infrastructure

Right? And I think at Cloud and Lithium you guys did an amazing job by basically with a small team being able to ingest majority of the web content. Right? And in my mind this is a gigantic achievement which many people don't know. Right? And don't realize that Cloud has all this technology to essentially crawl the web. So can you talk a little bit more about how that became possible? And like what is your scale? How much data you're actually ingesting? Yeah. So at this moment basically like Cloud is digesting all of the Twitter public data like the mentioned stream plus like all the data that users authorized us to collect like from Facebook, Google Plus, LinkedIn and basically some of the Wikipedia pages as well. And like this scale basically we process every day like about like six hundred fifty million new pieces of content. We were able to scale by basically efficiently implementing the collectors and then harvesters which normalize this data like in generic fashion. And then on the algorithmic side we were able to kind of hit a scale by being taking some trade-offs between like using NLP and using basically knowledge based dictionary

Mm-hmm. Knowledge based based dictionaries which kind of give you like much faster speed but you need to kind of cook them in advance. And yeah. Okay. So what is the balance? So you know like did the scientists say that it's 80% is data munging and data wrangling and 20% is science, right? How is your time balance? Like how much time do you do you spend on actually maintaining the systems and making sure they're up and running and performing? And how much time do you spend actually thinking of the science and kind of how do you prove the algorithm and how the algorithms perform? Yeah. So if you ask me it's mainly more like a 90-10 split. Mm-hmm. Because like for a good science like you need to have like a good engineering

Right. So kind of being able to build a system which is very systematic in its nature, right? Mm-hmm. Where you're not scaling only for the volume of the data but you're scaling for like a number of different data sources, right? Mm-hmm. You're having a number of different social networks and within each social network you're having different kind of sources you can pull from. Mm-hmm. Like from user actions to basically user graph or some other data sources, yeah. Mm-hmm. Okay

And so I mean every data team essentially has engineers and data scientists and there are different forms of collaboration. In some places scientists just prototype things in MATLAB or Python and then somebody has to take it to production. How do you guys manage this? Like how do your data scientists interact with engineers? Is there this difference or are they able to do different things? So the average profile in my group basically it's like a very like a strong engineers with like a research focus. Mm-hmm. So there is no like separation between engineer and the scientist, right? Mm-hmm. Because if you kind of can develop like the smartest algorithm in the world if it's very hard to get it into the production or if the time to get it into production is pretty long. Mm-hmm. And basically it doesn't cut it

This way you know like you engineer and you like do the research while doing it kind of you're all the time at the edge of the production. Mm-hmm. And you kind of push more iterate more and it just like works. At least it works for us. Mm-hmm. I mean I totally agree with this, right? So I think like you know if you're kind of good enough to implement the best algorithms it's probably learning software engineering practices and programming languages is probably an easier. Yeah. So you're kind of like a software thing, right? So but like you have to set up a process

Do you guys use continuous deployment, continuous delivery and testing? How do you make sure that your models are correct? Yeah. So most because like we use like a hive mainly, right? So that helps quite a lot because like the data catalogization is sold for you. Mm-hmm. Then a lot of you're forced to implement all of your kind of scientific algorithms within the UDFs. Mm-hmm. So basically your scientific chunks are almost like modularized by that. And then the data transformation is trivial because that's what kind of hive is. It's kind of SQL markup

So for us that's how it was able to scale basically. Right. But when you know when you develop new models and new algorithms, how do you make sure that they're correct? Like do you test your algorithms? Yeah. So basically like everyday like as things are running in production, you can very easily kind of like a build your model. It sends you, let's say like an email. Hey, your VCA file is ready. Mm-hmm. Or basically your model is ready and these are the weights and you can trigger like a kind of parallel branch with executes like similar, like a same thing that would execute in the production, but under like a dev namespace

Okay. So basically you're always able to run like, you know, like a five experiments and still be able to run like your kind of production daily. Okay. Interesting. Interesting. So, so then we've seen a great evolution of, you know, of cloud score and influence and topic understanding and you've been iterating on this. So, so where like, what are the next avenues? Like how do you see your algorithms improve? Where are your future directions? Yeah. So I think like on the topical side, like basically we saw like a problem of topical assignment and topical expertise

And so far we were like very heavy, like on the English side. Mm-hmm. And now like that cloud became like a part of the lithium, which is like international company. Mm-hmm. Like the next focus is basically being able to scale up for like the European languages, like in Arab, Arab languages. Interesting. How are you going to solve the, like you, you need to evaluate relevance and precision. Yeah

Yeah. How are you going to do that? So first by not committing to any kind of quality for the foreign languages, but joke on the side, like, because our approach is heavily based on the knowledge based dictionaries, like which are already kind of like international in their nature. Mm-hmm. Like we hope like to leverage that kind of to be able to scale it for the multilingual called corpora. Okay. And you know, I'm, you know, also running a soft scholar meet up and you guys are on GVM. So I am always interested in advancing data science on the GVM. So, and I wonder from your perspective, right, where do you see data science on GVM going? How can we make scientific computing on GVM easier? And how can we invite people who do R and Python, you know, as kind of GVM people, how can we make sure that it's the first class platform for scientific computing? Yeah

So I'm pretty biased about this. And if you ask me, I wouldn't even use the Java or GVM. Mm-hmm. Like I would use like C++ and it's something where actually you kind of have better control over the memory. Mm-hmm. And like it's faster processing and stuff like that. So probably not the best person. Okay

To talk about it. Cool. That's from the Google experience. Yeah. Yeah. They just open sourced MR4C, I think the framework for, oh yeah. But that's not the original Google framework. Well, and okay

And so, so, so basically you've been, you know, in this startup world for a few years. So how does that compare, like, you know, an ability to run this web scale frameworks at Lithium versus Google? I mean, can you speak about like, you know, does it become manageable, right? Can you do more with fewer people? Like what is the difference between these two settings? Yeah. I think definitely like it's easier to iterate faster just because like it's a smaller company, it's a smaller team. You're bounded by the resources, but sometimes that's just like the better thing because then you kind of have to think like in a different ways. A lot of, a lot of problems haven't been solved for you. Like there have been like in a Google, which has like a pretty good, like a good tool set and a lot of utilities that have been developed. So basically, yeah, it's more challenging, but again, like gives you more opportunity to innovate and kind of do things your own way basically. Cool

Well, thanks for sharing with us. And we're looking forward to your talk at Meetup tonight and the Tax by the Bay conference in April. 24 and 25 at Galvanize. Text.bythebay.io. We hope to see you guys there. Thank you. for having me here.