Devreal

Debora Donato: Serendipity Search, SF Scala @StumbleUpon 20150122

Debora Donato: Serendipity Search, SF Scala @StumbleUpon 20150122

Recording: Debora Donato: Serendipity Search, SF Scala @StumbleUpon 20150122

so I'm going to present a little bit stumbleupon the problem space what we're trying to do always histórico lee i can say historically as this company is a 14 year old we solve the problem what are the limitation of the current architecture and what we are planning moving forward and how this is relevant for scala ah you know passionate people so according to a study done from our marketing group there are over 30 trillion unique web pages this moment of the web and in average people just visit three or five different site each day so some upon trying to solve this problem try to take content and convey this content to people that is actually not searching for that particular content did I say that you have a really broad interest and you want to find something that can fit your interest but you are not having any keyword in mind so also upon trying to work in this space and trying to basically give you this personalized experience around some interest that you declare ah in general I'll do consumer find your content well there is a search engine ah you have a particular information need you died for some query and you retrieve a list of relevant result this discovery that is what we are trying to do ah you're not looking for anything thicker a seal you have some time to spend in a will say useful way and so you try to obtain some content from the web and there is the social feed to be facebook and be twitter or any other social network you usually use some weapon is not the search engine stumbleupon is not a social network even if people often I basically think that we work like a search engine or we work like a social net why it is different what are the challenges that we are trying to solve a stumble upon and that are not actually problem for search discovery is serendipitous again you do not have anything clear in your mind let's say that both Alexei and I follows cap Scala okay probably the content that is good for me is I don't know naive for Alexei and by 12 all right so even if we are looking forward to for same content what is good for F it's a totally different and nevertheless we expect that today discovery engine will be so smart to understand that there are different level of interest in a particular topic and convey the right content to the right people outside the search does not have this problem search is intent driven you know exactly or with a good approximation what you are looking for even say I'm looking for the address of a particular restaurant you can find it you cannot find it but you know how to evaluate the result in the search engine nozzle to evaluate the result discovery is more challenging because is once at a time so what's happened with discovery you just are exposed with particular piece of content there are two reaction you like it or you don't like it search engine up more than one chance to satisfy you because you are presented with the 10 let's say in general 10 result i come from research at yahoo for many many years yes there is a lot of investigation if the euro result is same first or second on third position but reality is that from a user perspective if that address is the third result that you find you are still satisfied discovery never repeat this what means once you're presented with a particular piece of content this piece of content will not show up in any future session this doesn't happen with search and in fact a search basically is use of bookmarking tool for many many user this is one of the top usage of search engine if you don't want to memorize a particular URL you memorize the query that allowed you to find that particular URL so it's a good market all this comedy adapt what does it means since we try to personalize the fee even if we are I don't know if you are looking for a particular topic and since we submit top content continuously day after day you will not have the same experience in two different day even if the user are similar this doesn't happen again for search engine result in the most of the case are fixed are there is a change that is less than ten percent in the vast majority of the result set of the query and last but I would say most important discovering this Taylor is trying to personalize and try to understand what you like what you do like search is still totally impersonal so how we do so we collect the number of signals that are implicit like the interests that you like could be dogs photography music but also we try to understand other implicit thing so for example say that we observe that usually you prefer trending content with respect to evergreen count something that we learn me while we are using oh we can actually use we have a mechanism of friendship that is similar to twitter and so we can present you with content that your friends like we will also collaborative filtering alike algorithm so if there is user similar to you we actually present you with the content that they like and we also use expert expert are people that seems to better spot good good content so they like a Content that the vast majority of the people that follow a particular category likes and we disassemble of signal to basically a recommend content how do we do this this is our to a lot of you know Iowa take whatever it takes to get the job done right and okay i know you are scala enthusiastic but i have met if this well but 14 years ago so but in 14 years what's happened is that the big elephant in the room is actually not a Duke but it's bhp at this point but we're moving so the amount of code you know that is written in PHP is decreasing the amount of code in Scala is increasing and we are much happier now than not a couple of years ago nevertheless this is the architecture overview so let's say the architecture that I inherited the six months ago when I join the personalization group so there are many different periods that allow us to do the job so in a nutshell new content needs to be ingested this happens with RSS feed so automatic way or through discovery by user enter in a Content analysis by planning which we do a number of filter bowler plate and some analysis kill destruction path to a sampling pipeline which we kind of try to understand how good this content will be how it going to perform once we will put definitely in front of the user after this pipeline the most of the content is throw away do not pass our threshold of quality and around five percent of our content is actually in Dixit and is a show to the user how we do this we have a lot and I will say a lot gain of offline machine learning models and they receive and whatever you have in mind yes we use them both that's basically every night all for training content news every hour do some magic sand computation this content is then third real-time through our recommendation engine multiplication elasticsearch there is a some part of our pipeline that actually use online events for some computation and true Kafka and this signal are basically used to create i would say day the final repository that is used runtime these are the basic component but there is much more these are all the other subsystem that actually helped us to at the job done and the lizard advertise service the you escort service that is very important for us that computed the a certain quality score of each page in our indexes and so far so how good we are in doing what we are doing these are our big numbers we have a three billion of total stumbled more or less per year we undred actually far more now our Android and Nikita million content in our index these are the list that people use to aggregate their content to retrieve the content and Mead and these are the amount of page that we add to our index per day after trying out a ninety-five percent of the content that arrived on the door Oh Sam papa okay so this is very very broad introduction ah it works we are in this business since 14 years so we people do not complain of the quality of the experience what we can do better okay and why we need the Scala for doing it better recommendation model how many of you are familiar with machine learning in general okay do you have any question for me I put there a bunch of key word right machine learning basic much learning basic course you know the one that you take on click on Coursera stem for the basic course so we use them all for it is on another we use them all there is a question that you want to ask me when you say this prediction you there is any question that you want to ask me well this kind of stuff is I will say it's very bad right ah well what what is available we don't use anything fancy but if you are passionate of machine learning you will notice that here that is nothing fancy what is matrix factorization what is it I didn't mention is not here if not here why we don't use it I mean you probably heard about the netflix competition 1 million dollar throw away to improve matic factorization basic algorithm what was three percent for purcell never use that we didn't even try it's not here why the problem is I mean in the machine learning space there are first problem and real problem having performant Algar it is not a real problem there are many other real problem for which we need better so multi factorization is biased it's biased again attempts at data tends to concentrate the most of the ratings right what does it mean this means that in Le you know business around recommendation think about Netflix there are a bunch of movies that collect the vast majority of items well since they collect the vast majority of items while you to have to use multi fertilization just go with them it works pretty well and multi fertilization will be bias against that either there is another problem in real space real recommendation on the one that you read about in search paper what is the real problem is the power law curse you know what is the power lockers the problem of course is that in real space when you have the user people the dev stuff that are distributed as a power law so there are few user that likes everything there are user did not do that not like everything in the search space is the same there are query you know what is the top we're in search at least when I was working at yo was Facebook can you explain me why the top query in a search engine is facebook i never understood user and i continue not to understand user but nevertheless million a million of queries every day from unique user about facebook okay we are here and we have exactly the same problem guys so the pump is that if you so when you do machine learning you need to train a model what you do you sample at random big mistake from user so if you get exactly the same power low so this is the head this is the tale that i announce it don't think about this being somewhere here right the problem is that in general when you start sampling a random on user just random as a point to you realize that you sample seventy-eight percent of the user but still who have in your user eaten matrix five percent of the ratings this mattress is extremely sparse what you want to do in this matrix nothing you know why this happened this happened because there are five percent of your user just five percent of user of that actually are responsible of seventy two percent of the ratings so you can do to think Oh random sampling user and so you will never get enough rating extremely sports or you get some poll ratings so you're sampling our sample from the hand of the power low I forgot you have you have that you're sampling from user they like everything what do you want to even think about create a prediction model for them they like everything they have the tendency of like everything so I'm just spending the time to do not nothing um when you start doing the job like you are you know dog a day at the college sample mirando take some lights put in your matrix factorization black box Oh usually in terms come to me and say the well i have ninety-seven percent accuracy yes but you're not predicting anything because this model is not going to work on real user um when you back getting in and smart way and you use the state-of-the-art Mattrick factorization la yes what'll it is you really couldn't see on this new product is this do you think i am going to production eyes this model no i mean sixty-three percent why should i it's a near a coin so what so what can help us to do a better job and actually enable machine learning in a meaningful way it's a question of more data no it's not that you we will convince user to rate more right what are you going to do paint them please freight now they are not going to do it a better model now why should and in some of machine learning algorithm Emma's netflix you know 1,000,000 model to work to improve my performance of what two percent three percent I'm not going to bother so if I improve the three percent by I passed from sixty-three percent to sixty-six percent I'm not going to production ideal so what we need the smartest data and all we get there we need to model we need to use machine learning to better understanding user before feeding some signal to event traumatic after iteration algorithm and when you see when it's a smartest what does it means you needed to understand your product you need to understand your user so 10 clocking here is added time spent from three category of user the one that raised the one that do all trades and the one that come down okay what is strike to me usually there is this concept of dual time the more you stay on the page of the more you like okay usually people that like like in the first few seconds that they see the page so the most of the like up and before you actually read the content instead of four dislike a page you actually need a little bit more time I think that the reasons because you need a kind of emotional reaction oh I'm really pissed off and you come down these are content that is not read this is the time spent all content that is not rated so are you so it is a signal and these are not five percent of your index five percent of your movie five percent of whatever you're selling if your Amazon or whatever it is these are all your stumble all my signals we have a model I don't want to to enter but we have a model that actually using please signal instead of writings using multi factorization on this model actually presented to an Akron sea of eighty-six percent from the original 63 on this kind of smart data I can actually try to do some improvement from the pure machine learning side but before I need to be here before you not to try to improve the performance of my algorithm training time is not a bottleneck I I went to a lot of jobs and a lot of which people say oh we can train your model in a cluster of thousand old in three minutes why should I care if my training takes three minutes or take 15 minutes I mean this is not really problem so what is moving forward so and where we need to modify the architecture to add this master data processing and here's the real performance bottleneck for stumbleupon for a lot of recommendation based company in the area so this is something that is a product that is currently ongoing and that is the creation of this pipeline that is based on implicit cinnamon bun operating our current data infrastructure ah characteristic of this i would say a dredge okay it's a few hundred of MapReduce jobs and workflow there are we have cluster on the round thousand of node we use scallop java scalding pig I whatever it takes to the gap and jump down fellas well over time I have to say some upon rely on this data pipeline but what if you want to use increases signal so what if you want to actually inject every single stumble as you know a feature to compute the quality of each page will talk about the real time data pipeline we are talking about better understanding trending URL that are Dells just detection just for a couple of hours so that are not captured by the nightly crop job they already know in in the decline how we do this so this smart data pipeline relying on a number of signal guitar aggregate signal think about variance or think about median over a stream of data how do we do this that the death the answer that's the answer and actually to tell you what we do this I will pass the microphone just for a couple of minutes to Sam Sam is actually building this ah and so as diverse area be we don't have a real socialist real-time data pipeline yet most of our data processing is done on a MapReduce jobs and other offline recommendation systems the first thing that we are trying to move to real time causing his real time you are scoring so because we are moving towards real-time event system we are also thinking about making the osmotic logic much more robust for example or like us what it's not just a one dimensional scalar quantity can be scoring by multiple buckets which would be based on your demographics or based on your device or a lot of these so we are coming up with the comprehensive topology a star topology that would increase the you know random availability of discourse and will also expand the scope of war between my source so the main infrastructure that we use are you are all the events that I talk here are consumed from Kosta and we use a combination of various Cassandra and a space to store the scores as well as user ID is for immediate cash so this would eliminate a big chunk or for offline recommendation and ETA five pence so the rough idea is something like this we have events coming from Casca and this is called apology is a very sunny Cassandra force in attaching and then we have some processed data from a historical data on HDFS and HBase final result and that is usually stored on an elastic search or a space itself so that's roughly what am I to do yeah oh yeah this basically conclude our presentation so there are this is just a use case we are actually moving to a different topology a lot of our sub component in this moment and we basically plan it to you know to take our PHP elephant out of the scene as up so we are looking for people 12 us to you know realize our vision using scala in line the next future so if you're interested because the address thank you doing the second equation ah whoa yeah yes sure when why ah good question quiet why so the question is why you're moving for PHP to scala ah again and the vast majority of our processing I was a sequential processing we have this code base that is actually was built over 14 years we have a code repository in which basically we don't share in this moment component we have what we call the big bowl of mud so reason is that when you start you know approaching the problem of computing some feature real time and you basically finally need to solve a problem like a thousand and thousand of signal that you need to see Beauty setting arrived on the door of your machine learning so basically the tool to get the job done are totally different so we we approach of company you know if color two years ago and we the problem of scalability how can we make this scalpel how we can actually have a signal that becomes part of our score in real time there is the water huge learning curve but derm asked these gentlemen set a screen before storm yeah so the idea is to move towards our tools for that allow us to scale in a better way and this is now PHP will be will be anachronistic right started doing something with a technology that is 12 year old