BDSBTB 2015: Debora Donato, Developing Big Data Components for a Recommender System
Recording: BDSBTB 2015: Debora Donato, Developing Big Data Components for a Recommender System
okay thank you i will start my presentation by thanking Alexei and the organized committee for invite me it's really privileged to open a keynote session speak before Michelson and Martino deskey and probably none of this person needs a presentation but everybody's asking why you are there and so let me spend just a little bit few minutes telling about my story in big data and what big data the different meaning of big data in my life as a Alexis told you I am a scientist I am a scientist wearing a hat and engineering at and what this means will become more clear me while I'm presenting my night of my life so my journey in big data starting 2001 probably am an old person in this room 2001 is when i start my PhD and is the year in which people start talking about web 2.0 and the same here in which wikipedia was born and none of the social networks we know them today we're actually there what data big data means for me what was the meaning well the meaning was 300 million node graph and 1.2 billion edges well it's not anything big in the mind of many people but it just didn't fit didn't fit my random there was no algorithm my language was C and I was implementing very complex data structure to implement external semi external algorithm to basically study the topology of the web that was big data in that year before 2006 what was my main problem at that point scalability no not really I just want my program to terminate and give me a result probably I wasn't one of the lucky person sometimes life is being at the right place in the right moment and i think i was in 2006 i joined your research soon after Doug cutting is when it is the year in which a dupe lend to yahoo and was a blast was really blast even if we know all the limitation that a due pass in this moment yoga research was granted first research purpose cluster initially 20 notes and after android notes what was really blast we were just doing massive data munging continuously machine learning was enabled we always were eating the the plateau of training curves on the majority of our problem were the year in which yahoo research was actually winning all the best paper award in many many conference and that all of this was enabled by by a dope but reality is that data mining and science were not done big data the vast majority of the training testing and product data little my pc is where were we doing the majority of my work and obviously using Harwick a python whatever you have earth and reality is is that big data didn't eat production so yes a dupe was powering the vast majority of click search but in a totally batch fashion and this is my opinion was far of being big data and 2012 when I finally left yeah what was frustrated data scientists were living in building f engineering was living in building g and we never met and we never science were actually ending in production in the way we actually think in 2012 i joined stumbleupon and the reason is i wanted to be in the area of data consuming company i wanted to to enable data sciences scale and i wouldn'ta to be able to eat production with machine learning and again i think i was in the right place in the right moment Kafka was released first release was 23 october two thousand twelve i do remember it's my birthday i was blessed best i will say birthday present ever and i was principal data scientist still 2014 what i was doing a lot of machine learning for generating of line rags for content understanding for user modeling bazzill this is not big data this is still small data and the vast majority of the company in the data consumer space that are data-driven are not able to be to use big data in a machine learning fashion in one year ago on the field suddenly i was promoted the senior director and the reason being is that I had a vision I had a vision that was not able to implement as principal data scientist and they wanted to implement and again I am a lucky person right moment in the right time spark and storm came to help me so what is my talk about you probably heard a lot of talking about making drug data drive driven decision at scale machine learning amazon service unleashed the power of your data but the reality is that is not always so easy with the same but it finally possible so i will present a little bit of the integration challenges that i am trying to solve in the last very last year and i will present to use case that goes beyond on analytics the vast majority of the toll cadet I year talk about analytic pipeline ah but machine learning is much more than analytic is much more than basic count and they really can help people to do a better job in many different applications and I will present to our the application where workings of so it's not my objective to convince you that the set of Scala based technology had become the de facto standard for big data I am a believer and if any of you are they as a doubt michelson and party notice key will take care of of it my problem is that even when you have a straightforward approach you saw yesterday you can implement and buy and the pipeline in one day well integration can become cumbersome when you have an ecosystem that is 14 year olds that does not respect any of the modern architecture paradigm and needs to change meanwhile you are trying to implement a total new technology but it's still possible so we want we wanted to try to moving from previous language like PHP to scala in a with a seamless integration so what is your problem legacy what is legacy may experience legacy is an ounce of cart it's really unstable it's something that probably over 14 year you did something pretty incredible with it and you are there and you are looking at let us say wow this stuff it's really cool I really can do a really incredible thing but the only thing that you don't want to do with it is touch it don't even try I try for one month and I said okay forget it let's let's do something else this is what I inherited one year ago I say okay simple simple structure right just for the people that doesn't know what sample ponders you register you declare your interest that could be computer science machine learning get cooking whatever it is and after this we start the routing content to you we learn from the interaction that you have with with the search engine and we tailor the flow the information so you don't need to search for information the information is routed automatically to you we work in a i would say between entertainment and information we made of information a form of entertainment so as i told you this is 14 year old ah was written total in PHP i say totally but sometimes you find some peace in java some peace in PHP some pearl in pearl that is kind of weird and the problem with this data structure is that there are a lot of component on functionality that in the model time you probably would have implemented standalone microservices that if that are embedded are embedded all along your monolithic structure end to end so how you do that so when the majority of your small team and some weapon has 44 engineers in total right and my team at that manage all these is eight people the correspondent t teams in Pinterest are like six teams so just to say to tell you so when you're the majority of engineering is working on keep the legacy running right and modify the legacy Ronnie when your data scientist now I just use our no no I just use Ivan peak and they don't want to emigrate and when you want to innovate what are the things that you care first of all you need to leverage code that is written in many different programming language you cannot implement everything from scratch everything is very complex you want to have the ability to receive data from a large variety of sources I don't say anything new when I say that I want something that is fast that is fault tolerance and can process events in an incremental fashion and in most of the case you want to have a purpose-built cluster you don't want to mix up with legacy oh wait arbor that too old and you want to the Commission what is the choice in this case so the first thing that I decided to to change in in this in this monolithic system was ingesting pipeline currently I when Justin pipeline takes n2m to end around 10 minutes in this 10 minutes we had to do an incredible amount of think we don't work on catalogues Netflix or Amazon or all the big rock engine that you usually deal with F catalogues so they exactly know what the content is about someone we don't everything that is out in the web trillion Oh Paige can become a candidate for recommendation user come to us they ingest the content and we don't know anything about the content ah so in this ten minutes we try to understand everything we try to understand the language we try to understand the content we try to understand the format is a news is evergreen is a video whatever it is and we do some sort of classification which interest match is a computer science is machine learning and even if it's machine learning is machine learning for a new Miss with machine learning for a data scientist the principle that a scientist currently it takes ten minutes so this awful this awful right ten minutes what if you want to do anything if a user coming just some form of content and you want to leverage this create an engagement experience you want to say okay you just submit a content on machine learning or scholar I want to react and present you what similar I have in my repository for you current methodology doesn't allow how do you solve the problem well you want to find similar items you want to find items that dig further in a specific topic or you want to dig further in brother topics and also you have some form of constraints right you want the very quick injection to recommendation turnaround time around ten seconds ten minutes right now good luck with it you need to adopt streaming processes with at least once guarantee I cannot lose any of the content that is ingest I needed to build it important system so i don't want once something is processed I don't want to reprocess ever again and I want to capitalize on non-linearity when I have there the possibility of doing that and obviously for retrieval of the rag I need low latency so I needed to pre-compute rags and I want to retrieve rags in constant time and obviously horizontal scalable design but this is not news so what was the scientific approach behind it well first of all you do your batch work right you select a I quality data set for your training and for your testing and you do this and online for each URL ingested you want to extract all the text feature that characterize URL you want to compute topic caching filtering a noise keyword finding generic topics specific topics and computer similarity relevance without the pipeline look like it's a pretty straightforward pipe allied I believe you want to pursue you want to detect remove the boilerplate detect the language do some sort of cleanup you'd never have a page with one language do some non chunking use whatever tool you can use and this came with legacy wikipedia notation but you're learning NLP tools at this point you want to call a stag new york time and nyt are the same and you need to know and you want to compute the tag score the technology we use is pretty straightforward right you want to present a similar document to you what you do you map your document in a complex topic space and you basically compute a similarity measure between this topic space are you all familiar with lda with familiar not not many okay so lda is it supported algorithm in as an unsupervised algorithm in towards what we try to do is observing the world in each in each document we want to understand what are the topics that this document is about so if I talk about altiris sheep obviously my document is not just about cooking but is also also about health but if you are talking about and nutella cake chocolate whatever it is there is no alt in this in this document right so as you can understand a lot of document can be interpreted look a different you know dimension this tags this dimensional topic space is when you want to work this is an unsupervised method and it's pretty standard and basically what it does is generate the topics automatically like this that you see is pretty meaningful and you can map each of your document in the space ah what was our choice for implement all of this pretty straightforward right we obviously obviously decided to use a storm storm allowed us to basically respect all the service layer agreement that we put in place so we use previous legacy component and we wrap in order to send events to the CAF Gibraltar and we implement a classical storm topology written in Scala and so I basically force my engineer to start using skala and we basically generate topics and we save in HBase reason being is that we needed for batch computation every night to regenerate topic on the fly and something that I want you two to notice is that the for the topic model we actually use of data why why using data because machine learning is still small so even in this very complex environment ah actually we can train and test our model on a single machine twenty four cores um and this is an important aspect the spark and all the days callobus technology have enabled right machine learning a scale but reality is that Dee's happened one year ago the vast majority of data driven company or that a consumer company are still collecting and in fruit you unleashed the power of data on company that never use data so far and the problem that they are trying to solve a really simple problem and they don't live in the complex phase your dimension of space is small enough to still fit ah one single machine this is what the pipeline is about a little bit of numbers but the important thing for me is that yes we got it this pipeline now is able to process URL in 10 seconds so 5 in 5 seconds no no no 10 seconds I'm joking and I want to present another case ah the other thing that I start that I work I wanted to work well what we call us quality score so think about user rating your content right ah in the most of the case basically all the company and you can totally see this when for example you stumble upon pinterest do not react runtime of ratings there is a bad job that you're in the night do something and the day after you see the effect of the user behavior on the recommendation engine so you need till at least 24 hour but this is crazy right I want to react in real time for the user and their content like news or like a trending content for which I won't understand what's going on meanwhile is happening I don't want to wait 24 hour to understand that that was the news oh thanks god it's gone the problem is that the very few user give ratings in general is less than five percent of all the interaction in any social network Netflix stumbleupon pinterest you call one less than five percent of iteration are actually light so you don't tell information about that iteration but you can use some signal you can use a time spent a number of clicks if you scroll the page assembled upon we just get the thump time spent right so what we thought to do is this okay let's take this in real time and let's take the information that I collect from each user so far and in real time meanwhile is happening next tumble I decide if the time spent given a number of feature mean that you lie to you dislike the page we implemented this pipeline and we got to seven seven percent of a currency this is a great result it's a great result because if you think of the netflix competition challenge are you familiar with the neck put netflix competition challenge so they had exactly the same problem five percent of ratings same density of the data set that i tackle every day they spent 1 million dollar right I don't have it 1 million dollar and even if I a bit with what with this kind of data implementing increasing ten percent of the currency will brought me to an accuracy of seventy two percent that I still doesn't want to use in production using this kind of algorithm I can actually bring the density of my recommendation matrix from five percent to thirty five percent and multi factorization with these data touch an occurrence eve of eighty-three percent so I can actually go in production without spending 1 million dollar this is what we implemented and the problem that I had this time is that I didn't know what I was doing what I mean with this this was a research problem I didn't know that the feature that I select was the right one I didn't know if I could do better I didn't know if the machine learning algorithm that I was choosing was actually the right one so I want a real time machine learning pipeline for research purpose I want to be able to change my future without redeploying everything I wanted to be able to change my learning although without redeploying everything once again the two witches storm for this everything is implemented in skala ah and basically this pipeline nice is fully integrated with the legacy i'm concluding my talk saying something that so it's exam I'm climbing here that there is no need for big data machine learning no I'm not saying this I'm saying that when you start in a legacy there are a lot of constraint but i will say that a petite comes with eating so after we implement this to pipeline and we actually integrated successful in our pipeline we said okay now we want to tackle real complex problem and now we want to learn a number of dimensions are data stream large and we eat the limit this is actually the project that we're working we are trying not only to model a content in a tag space we are trying to model the user in a tech space at the same time and we try to learn all this model so in this model the only thing that you know all this thing in the square all the things that you see in the circle you don't know you want to learn when you actually do these even implementing on parallel structure not using I would say the right technology training could take days so when you are in the situation and we are in the situation you actually eat the bottleneck in batch so what is next you cannot do anything at this point if you don't use scala if you don't use spark so and i think that the message that I want to convey to everybody is that when you think about machine learning a scale don't believe to buzzword scale your system think about what machine learning means for you think about the complexity of the problem that you have think about our integration play in everything on whatever you want to achieve but when you need use spark and your scale thank you