Devreal

sfspark.org: Deepak Thakral, Vivek Hariharan, AdTech Engineering at Chartboost

sfspark.org: Deepak Thakral, Vivek Hariharan, AdTech Engineering at Chartboost

Recording: sfspark.org: Deepak Thakral, Vivek Hariharan, AdTech Engineering at Chartboost

my name is Deepak Cockrell and we have some very interesting sessions lined up for for you guys we're going to kick off by just giving you a quick introduction about advertising and how the advertising system ecosystem works then we are going to talk a little bit more about who chart boosters and how does ad tech echad boost work and then I'll have my colleague will rake come and talk he'll give you a lot more in-depth information about chart boost and about our ad tech ecosystem about how we utilize spark how we utilize Scala so there's a lot of interesting depth and tech deep dive into each of these topics and then we leave some time for Q&A and then we will have Alex from Cadell which is another advertising technology company and he's going to give another technical session about how they utilize some of the technology components at retail right so Oh what happened to the go blank okay things are good so first just for a quick show of hands how many of you have either worked in advertising technology or a familiar with it yes quite a few okay how many of you use spark okay good so with that let's let's kick it off so first of all mobile advertising why it matters right we're going to talk about that packet scale how our decisions made in like 100 to 200 milliseconds the time you take to blink your eye there are decisions made about who you are what add to serve all of that happens so we'll give you a quick kind of 3d of that then I want it off my chart boost Google will help give you context when the big comes in and talks about how we have designed our advertising infrastructure here at Rogers and then we'll talk about Scala and spark what are the use cases how do we utilize the technology here and then we leave time for Q&A okay so I wanted to first start off by you know talking to you about how did advertising evolve right there used to be a time and advertising a lot more simpler and I actually have this ad from the 1960s which is you'll see very shortly okay what this ad is about I thought I've started off with something a little bit more fun and light so take a look I think we need to get the volume up today my skills come [Music] so this was an actual ad from the 1960s well you could get away with whatever you wanted to say that it actually like has food flavor you know it reduces calories it makes you thin right but like you know advertising is evolved since then it's become a lot more sophisticated and the format's have changed they've gone from live TV to desktop to mobile web to mobile in-app right and cook still continues to advertise except that is just sorry let me get this out of the way coke still continues to advertise but now the format's has changed right and what you see here is Selena Gomez on snapchat right where the ad is most likely like a vertical video which is like five seconds long but it appeals to the audience and then you look at on the right-hand side as a Facebook ad which is like a carousel which shows different coke flavors right so the brand of Coca Cola has not really changed but it has evolved and so has advertising what's more interesting is that the technology that powers this has gone a really long way so what do we see in terms of Technology is that brand and performance are starting to converge right and what I mean by that is that there was a time when brands just care about putting their ads on TV and reaching an audience they didn't quite know how many of their audience actually cares about the ads how many of their interacting with the product but today what happens is that people have so much of interactivity the smartphone women on their hands that they can click into it they can interact with it so as a result of that there's a lot more like information about how many people are viewing the ads so in terms of viewability what calls to action are being made right what interactions are happening what percentage of people watch 25% of the video have to video the entire length of the video right this is there's a lot more analytics coming in and similarly you know this information that's happening around bidding so basically once we know that there is an audience or user in front of a phone there is a lot more information available such as what kind of platform is the user on are they a male are they a female which country are they and what are their interests what kind of demographics have been and all that information is utilized and within like sub seconds decisions are made about which is the right ad based on the targeting parameters to show back to the user right so there's also information about interests like what kind of interest do they have do they like to play soccer would like to play baseball do they like certain kinds of goods and again decisions are made based on their interests similarly you know we create things like look-alike segments which is once we identify that if you're a user who likes X then we will utilize machine learning and artificial intelligence to understand what are some other users who also like X and therefore make and show ads based on that right so what you're seeing is that there are certain things like behavior targeting and I'll talk a little bit more about how chartboost uses behavioral targeting because we have a lot of information about gaining audiences when we are able to utilize that and say if you happen to like casino games now let me show you ads where you're going to see more casino ads right so the point I'm trying to say is that there's a lot of data that's happening behind the scenes and therefore essentially a lot of advertising companies are in essence big data companies so they harness all this data and this is where technologies such as spark come into play because they help you utilize a lot more machine learning and cloud computing to understand what kind of ads to show to the right person so why does mobile advertising matter it matters because there's a lot of dollars going into it as of 2017 is projected that we expect to have about 140 billion dollars flowing globally into mobile advertising and that number only keeps going up right so we expect that by 2020 it will be close to 250 billion dollars right and therefore the next question you might ask is who dominates this mobile advertising and what's going on behind it so what we see here is that three formats dominate mobile advertising they are search social and display search you might be more familiar with if you're on your phones you go to Google it's the dominant form of search advertising right social we see very clearly that there are a couple of social platforms that dominate mobile advertising but that space the Google or you know Twitter or snapchat but the display ecosystem is a lot more fragmented and we'll show you how display advertising works right in terms of you know that's kind of the heart of powering ad tech and how we have this entire ecosystem of demand side platform supply side platforms and how they are interacting in a real-time auction bidding ecosystem to decide which is the right ad to serve to the to the user in that publisher so we talked about social right so we definitely have Google participating in a lot of ways but we also have the Giants with Facebook Twitter and snapchat and they make up pretty much the dominant form of advertising today in fact in terms of social advertising what we also have is what we call as the programmatic ecosystem and what you see here is basically audience which is users right and so these users would be in front of different types of web sites or different types of mobile sites like they could be in front of apps so they could be in front of cnn.com or the ABC app but what's happening behind the scenes is that when a user is on this web site or on the mobile app a request is being made to an ad server the ad server will then have the ability to run an auction and see that do they have their own set of advertisers if they don't they will send out a request to an exchange and that exchange works very similar to what you would think of as a like let's say like a stock market like like the New York Stock Exchange where there's automated buying and selling of stocks happening in the advertising world and Ad Exchange is facilitating buyers and sellers and deciding how you can automate the buying and selling of ads so all of these ad exchanges like double celaya cap X's you'll be calling etc are going to then send their ad requests to multiple demand-side platforms and these demand-side platforms will then make decisions based on the machine learning and the Big Data capabilities that they have as to which are the right campaigns and pretty much in a split second will decide that do they want to participate in that auction what should the bidding price be and how much are they willing to bid to acquire and show an ad to that user and then from there on you know the DSP makes the decision it represents that advertiser and then returns that ad response back and all of this happens in like about 200 milliseconds so either rather than me just talking about it will actually show you a quick video which shows how the programmatic ecosystem works advertising technology in those the target Django if there is against a magic changes profile and a desert if no campaign targets change though the server sees the mation impression programmatically requesting responses selective traders and networks and supplies nine platforms or SS machines impression is not cleared to super-amazing declare the impression in a programmatic directly via private exchanges in the impression is again not cleared the request is sent to an open at exchange it hosts the cheating including open that exchanges a big request containing information on jane doe's browser website URL and had typed multiple videos including traders and networks and demand-side platforms or DSP each bidder processes big requests overlays it with additional user data and marketers targeting and budget rules traders health evaluates the request select the create and send it along with optimal bid price to add exchange ad exchange selects winning bid from readers responses per second price auction ad exchange sense of any cash you are elton price for watching pushers answer publishers ad server tells jane doe's browser which had display jane doe's browser pulse and winning leaders handsome and sense matching hands browser browser displays web page including matching ad winning bidders ad server receives ad tag data munging does condition interaction experience and that is 200 milliseconds the life of a programmatic a depression for more information visit us and crossing calm so you could see right there right like in 200 milliseconds how many different computations are being made how much does that add requests flowing through to different entities like an ad exchange to multiple demand side platform and the auction is happening and then the response comes back and you as a user you're just like you know scanning the news or whatever and you see an ad but you don't realize how much complexity is going behind the scenes to determine what the winning ad should be that should be you right so with that what I'd like to do is shift gears a little bit and tell you a little bit more about chartboost and how do we you know utilize advertising and a little bit similar to what you saw before in the ecosystem so as a company we're back by Sequoia and TransLink we have about 100 people between San Francisco and Amsterdam and as a platform that really helps game developers monetize we see somewhere between 700 million to a billion unique monthly devices each month right and what that does is that basically we we help app developers make money especially like we're the world's largest platform for monetization for gaming audiences how do we do that we have our SDK and over 300,000 apps so basically 90% of the top grossing iOS and Android apps in the gaming ecosystem utilize our SDK and with that we also have large like gaming Studios like EA a Zynga you know Big Fish King all of them utilize our network to basically drive user acquisition and so because of this STK reach we see between 700 million to a billion unique users it allows us to create a lot of rich personas about their gaming behavior so we understand what kind of users tend to like casino apps or what kind of people like like trivia apps or casual action and then based on that we are able to build these kind of rich personas right or grass so where we going to spend more time talking to you about how do we create those grass how do we make the right determination about which ad to show so that's a little bit about us in terms of our SDK penetration we're like number two behind Google AdMob in terms of the SDK and how many users how many developers users not just at the head but also the torso in the tail so our a wide reach of in terms of our STK penetration then the question you might ask is how does advertising work for a mobile ad network to simplify it it's not just our boost but any other mobile ad network so imagine that there's a user called Joe and Joe loves to play subway surfer so how many of you play mobile games by the way Joe hands maybe a few right so if you're a mobile gamer you know what you're going to do is you're completely immersed inside this game and when once you reach a level what we decide is that it might be the right time to show you an ad or you may also decide that you want to earn some points to continue into the next level so when you want to earn some coins you have a choice you could watch a video and that video basically tends to be a promoted video which is coming from an ad network so at that time the mobile game publisher will make a request to the ad network in saying hey do you have an ad and so what we do is we utilize our SDK to then show different ads and so what you will see is for example the bubble Witch Saga which is another game that will decide to show as an ad or as a video so what are the different types of ad formats that we have we basically have what we call it's playable which are more like interactive games they're like mini games that are built into a utilising html5 into the ad unit so that it's almost like a try it before you buy it the user can decide that they want to interact and play the game and if they like it then they'll go to the App Store install it another example of videos where people get to basically watch a 10 or 15 second video they tend to earn some coins but they like as they watch the video they tend to like the game and say oh yeah this is something I'm interested in let me install it right and then the other one is just say basically a static banner ad so what happens is that when people click and go to the App Store install as 100% performance driven ad network we only get paid when the user installs the app and opens it up right and so therefore there's a lot of science that has to go in for us to understand what kind of users are interacting with these ads what kind of users are actually installing and opening the ad and how do we create more installs that generate more revenue from us on advertising technology perspective so what's happening behind the scenes what's happening is that all of this information about the user is going through the chat Boost SDK so information like what kind of platform is it is it an iPhone what kind of OS is it is that an iOS then user is this person which country is this person is in the United Kingdom what game are they playing what is their ID FA or their unique kind of advertising ID and then we're taking in all this information and then it goes to the ad server the ad server doesn't say how many advertisers do I have who are eligible to be considered to show an ad right so we'll make that determination then we'll make a request to add relevance module which will then make a determination that off the hundreds of advertisers that may be eligible to serve an ad which is the ad which has the highest probability of being clicked upon being installed upon and we use something called as eCPM so we'll come up with a predictive eCPM and then based on that highest score will then decide which ad to send once that ad gets sent and rendered then the user may decide to interact with it if they click and they go to the App Store and install we have something called as attribution by which we actually get to know that install happened and so all of these billions of events that are happening on a daily basis in terms of impressions clicks installs and even post install metrics like how many purchases are being made or what kind of users are coming back to the app on day one or day seven or day fourteen all of that information is being collected into a data warehouse and being used to compute which is the next ad to show to that user or to another user and with that you know hopefully you get a better sense of kind of where the advertising ecosystem is we showed you ads from like the 60s to now we kind of talked about how programmatic is evolved where within a you know 200 milliseconds were able to make decisions about ads and then we talked a little bit about chartboost and the mobile advertising ecosystem so I'm sure you interesting to learn about how you utilize some of these technologies in or in our technology stack and for that I'd love to have they come and talk about you know do a more deep dive thank you my name is Vivek and I manage the engineering team at chartboost and what I was saying was that I really need that cook from 1968 that makes me leaner I mean pretty interesting at music deep deepest walked us through the chat Boost Mobile mobile ecosystem and what happens in the chat of the chat was sgk but now what we would do is go a little bit more into the detail of the technology behind it what happens behind the scenes typical our prior to joining charge was I was with with Yahoo for 15 years mostly working in at Tech and what I've seen is that most ad systems have three big components one is a web application where marketers engage with the web application to build campaigns and creatives then that is like a serving cloud which is responsible for serving ads and then the third is the data cloud which is responsible for running analytics and machine learning and whatnot let's look into the chart boost high level architecture what happens like just before this slide deepak showed us what happens between and the user playing a mobile device to an chartboost sdk to an ad server serving an ad now let's little bit dig into what happens behind the scenes inside the chartboost garden to begin with and I'm a marketer wants to promote his newly built apps he is engaging with this web application and is creating an advertising campaign on the other side a publisher who wants to monetize money monetize and make money he is going to create a publishing campaign using this web application the publisher the addition as part of creating the publishing campaign is going to specify an ad location what kind of ads need to be shown inside his publishing app on the other hand the advertiser is talking about what kind of users you want to reach how much he wants to pay for every install and what's this budget all this information are getting into a database we use a chartboost from where ad server sucks it into in memory so when a user or a gamer is playing a game which has the chartboost SDK installed in it that SDK is making a request to that server that ad server is going to perform a bunch of business rules a bunch of validations and determine the best performing ad that need to be returned for that ad request you go a little bit more deeper into what the best means how the ad server determines what's the best ad is in the subsequent slides later from that on the ad server is also emitting a bunch of events such data into a cop car here we use Kafka as a message queuing system which are partitioned by date all the data are getting into the cop car then later sucked into a data warehouse we use head GFS and high as our data Mart right are inside the data line there's a bunch of aggregation pipelines are running bunch of machine learning models are running a bunch of streaming pipelines are running to perform different use case so again we'll talk a little bit more about those in the subsequent slides and then later on a bunch of repose aggregation metrics are generated which you are which are presented to the advertiser and publisher through different side of the foot okay let's slow double-click little bit into what's going on inside ad server and what's going on inside the data the data slash modeling world so inside ad server there are two main parts inside ad server then the top path is the get get ad path and the bottom part is the design spot right so let's talk a little bit more about the get ad part that when the SDK makes a request to get an ad the chartboost ad server can do one of the two things it will also it will first try to find an advertising campaign that chartboost has chat boosts knows about and see if that ad is a good ad that can be returned for that ad request simultaneously is also submitting that requires to a chartboost exchange chartboost exchange follows open our TV and sends that request to a bunch of integrated demand-side platforms examples include Rubicon add colony and more these DSPs may have some advertising campaigns from advertisers that are not directly watching the chat boost it may be an ad from Coke or a Pepsi that are not working with chartboost yet and maybe some of these advertisers may be interested in some of our users and in case they are interested these DSPs are going to submit a bed for those requests into the chart whose exchange charge boosts exchange and chartboost ad server are going to determine which is the ad that is going to perform the most money for the publisher right and these DSPs compete among themselves as well as with the chat boost network and the best performing ad is going to be returned to the publisher so that's what happens our to the SDK which gets shown in the mobile in the gaming app so that's what happens in the get Ad part in the event spot sorry before going into the event spot a couple of additional things that goes on in the get pad typical problems or use cases that are addressed in any an attack or ad server are frequency capping and negative so most marketers may want to control the number of times an ad is shown to a device as well as they may not want to show an ad to a device if the device has already installed that apps we call a negative targeting so to support those use cases what we do is that we load the last seven days of impression data that is the data about what are the ads that the device had seen in the last seven days as well as the click data and the install data into a key value store superfast key value store we use aerospike for that in the events path what happens is that typically unlike the web world an ad is not immediately shown as soon as it is returned to the publishing app the SDK may choose to cache the ad and decide to show at a later point of time as decided by the gaming developer so when ID is shown the SDK is emitting a beacon back to the ad server and similarly when I ad is clicked or installs a beacon from an SDK is being sent back to the ad server and the ad server is emitting all those data again back into cough cough so that's this bottom part the more interesting thing that is specific to mobile in app ecosystem is attribution what happens is that unlike the web world in the mobile world people pay for a particularly in the mobile in app world people say pay for installs and that is a lot of that is most of the time there is a gap between the click and install what I mean by gap user may see an eyes may click on an ad and may good the user is taken to an app store in an app store he may choose to download the app and there by the time the app gets downloaded the user has gone to do something else the user has not yet open the installed app right so from the time the user clicks to the time the user downloads and open the downloaded app there is a delay and during the point of delay the same ad might be shown to the same user from a different ad network while the user is in an another experience right so now there is a conflict as to which ad network gets credit for that install and both the multiple ad networks are saying that I showed the ad to this user I need to get paid for this install so typically how this problem is solved in the mobile ecosystem is by using a third-party attribution provider some of the examples include tune and adjust what happens is that every ad network sends a event to the attribution partner every time and user clicks on an ad and then the advertiser the under advertisers themselves sends an event to the attribution partner when the user installs the ad now this attribution partner needs to do the job of arbitration who gets to get the credit for install and typically and there is a variety of attribution methods being used but that commotion common one is to use last click attribution what I mean by last-click attribution is that which ad Edward gave the most recent are the last click just before the install happened and installed here means the user opening up the downloaded app and playing the game that's what we consider an install event so that's what happened here moving on let's talk a little bit more about what happened in the data slash modeling world so we saw about how an ad is served by the ad server we saw about how all the events are getting logged into the Kafka and from Kafka the data is getting sucked into a high such HDFS data mark typically there are if I want to divide them pipelines that are running into in the Hadoop cluster into four categories I can divide the pipelines into four categories there is a set of pipelines that are running that are generating metrics that is accusations by for like what is the impression count what is the click card what is the install count by different dimensions and some of these metrics are generated real time by spark streaming applications and some of them are generated in batch mode using hive an air flow so that's one common use case the second use case that we do is the budgeting pipeline what happens here is that we have a real time streaming pipe which reads the events money our transaction events and aggregate those transactions by the advertising campaign if we want to know how much money an advertising campaign is spent right and those aggregations are written back to a new Casca topic the data that is written into the new Casta topic will typically look like advertising campaign ad and spend for today that data is read back by the ad server to use to an ad server uses that information to determine when to stop a campaign as it reaches close to its daily budget limits so that's the third part I've shown in the diagram the couple of the other popular use cases that are going on in the data line are machine learning and audience targeting the the second path that we are showing here is a machine learning path where we read impressions clicks and installs over the last thirty days from a spark pipeline read and train a logistic regression machine learning water and the job of the logistic regression machine learning we probably go a little bit more deeper into the subsequent slides but it is used to compute a model that is in turn used by the ad server to compute install probability or predicted in CPM for different advertising campaigns given an ad request and we use that to rank the campaigns and compute the best performing at the top pipeline shown here is the audience targeting pipeline what we do here is have a real time spark streaming pipe that leads boot up events want to spend a little bit of time explaining what a boot up event here means in the mobile world a boot up is an event that the SDK in it every time a user opens up a game gaming app so every time I user opens up a gaming app that boot up event generated is read by a spark streaming pipe and return to a key value store we use dynamo here here the Dynamo is indexed by device ID and for every device ID for every app we store a bunch of a data points like the boot up count in-app purchases the last boot up count the install date and we use these informations to do a variety of interesting things let's double click little bit more into the machine learning pipeline and the audience targeting pipeline okay so first we begin with the machine learning pipeline so the goal of the ad selection is to choose the best ad given an ad request the keyword used here is best so how do we know which is the best ad to show so let's step back a little bit what is charge booth ad server is trying to do right it is on one hand we have publishers whose goal is to maximize revenue they are there they have already a game which is having a large set of users and they want to make as much money as possible which is measured in terms of eCPM our money made per thousand impressions right on the other hand we have advertisers who are releasing a new gaming app and are interested in getting more and more installs at a lower price point and not only that they are also interested to maximize ROI and ROI here means the quality of the install they are not interested in getting installs where a user comes installs an app plays for a few minutes and then quits the app and never returns back they don't want such install instead they are interested in paying more for the installs where the user downloads the app and engages deeply with that app and makes a ton of in-app purchases so we have pretty different goals for a publisher and a pretty different goals for advertisers and a chartboost in the middle is trying to match and align these goals so before going deeper let's double-click a little bit more about what is eCPM so eCPM which means expected revenue per or cost per thousand impressions I say expected because it is the probability that this ad is going to make money in future over the next thousand impressions so it's a predicted value or not the real value that's computed by multiplying the install percentage it them their money per install our bread per install with thousand so in this equation both the the thousand multiplier and the bed are constant what I mean by that is that the advertiser as part of building a campaign chooses a bit value for every install that the campaign get that's not changing and the multi-thousand multiplier is also not changing what's changing is the install probability which is what we are going to compute using our using variety of algorithms so how can we compute this install probability so in a in a very very simple world what will happen is that you have here two advertising campaigns at the top we have a bunny pop rewarded playable campaign and in the second row has a koa PGT video campaign the top one has an install probability of 0.5 person while the second one has an install probability of 0.1 percent however the bunny pop advertiser is paying a CPI of our dollar 25 cents while the KOA is paying a CPI of $9.50 so when we use that formula eCPM formula we get a predicted eCPM of six dollars for bunny pop while we get a predicted probability of 14 dollars for koa so here the second advertiser wins because he has a better predicted eCPM so this is a very simplistic world but the real world is not more complex what happens so what I mean by that is that this is looking at the advertising campaign across the network but in real world what happens is that the performance of the advertising campaign varies significantly by different dimensions a particular advertiser will perform really well for our particular publishing app or for a particular set of OS or in a particular set of Geo countries and whatnot so here I am going into an example of the performance of the stay advertising wrap by different publishing app so we are taking racing penguin as the publishing app and we are taking the same to advertise the bunny pop and koa in this case because the install probability is now in this racing penguin publishing app is so different by using the same formula to complete particular eCPM here the bunny pop is performing significantly better so then we looked at the whole network the KOAT advertiser was performing much well but when I looked at the racing penguin publishing app actually bunny pop is performing much better now so in this case bunny pop is the winner now if I extend this with a variety of features our variables country publishing app publishing category user preference connection speed creative ad format device model platform and many many more right each of this variable is important right when I expand all these things quickly the function to compute the install probability by these variables becomes complex and this is where we use machine learning so we how do we do that so we take the historic impression data and extract variables we use the spark a Scala code to extract the variables the variables here include advertising campaign publishing app ad type country right and the trainer model what do we mean by trainer model basically here a training a model means that we use an algorithm in this case logistic regression to compute them or learn a mapping between these variables to install probabilities so that's what we mean here once we determine a model what we are going to do is that we are going to use the same model and the same feature extraction code inside ad server to score the advertising campaigns let me spend a little bit of time on that what that means given an ad request that could be a variety of eligible campaigns now we need to determine which campaign to be returned we take each advertising campaign and then extract the different variables for that ad requests like publishing AB ad type and country and use the model where the model has a set of weights for these different features and compute the install probability and from install probability we compute the predicted in CPM and we use that to rank all this advertising campaign and return the best performing advertising campaign okay some real was learning so we had a very interesting journey the first mistake there are a mistake and a power that spar gives us this part provides view distributed machine learning ability but that what happens is that you get a power to run machine learning at a very very large scale so we started our journey by training a model with advertiser publisher combo which quickly generated into some 30 million features we trained a model where we hashed down 30 million features into 1.5 million feature buckets and trained the spark model for those 1.5 million feature hatch buckets quickly the model became too complex it was very hard to troubleshoot figure out what was going on we are not able to trouble du Quai efficient deep types and analyze how the coefficients we are looking for different features but additionally we were also using different models to Train different kinds of campaigns right like we were like the really big top advertising campaigns were trained using a specific model while some other smaller advertising campaigns were trained using a different machine learning model and what that brings is that the different models have different variants and different eugen right I'm so to combine the you CPM scores from all these models and rank becomes and became a nightmare so what we did be simplified simplicity is the key then even though you have a lot of power with the distributed machine learning simplify things so what we did was that first step we did was that we used a technique called frequency filtering where we looked only at the features that have seen a enough number of times in the training data and crashed all the other features that did not that we are not seeing enough and additionally we used we we categorized the publishing app into a big publishing app and small publishing app and use a specific one particular modeling function to score all the advertising campaign for an ad request this avoided mixing models per request and greatly simplifies the model we also lifted feature Aisling which enabled us to do coefficient analysis and we saw a significant lift in the model performance before going into the second use case want to say one last word here so we spoke a lot about the model models for predict user ipm prediction we also are working on ROI and LTV predictions right advertisers care about not just installs but advertisers care about high quality installs like they care about purchases acquiring those purchases they willing to pay more for those purchases so doing Auto analysis so we want to have baked into our model so that we can we can predict the purchase we can predict the users who are willing to who are going to make lots of purchases in the game and charge a different key for those users with the advertiser so that's a pretty complex topic it's meant to be explained in a separate different session so I am NOT going into that in this today's session but I just want to say that that's a fairly complex machine learning topic and we are working on that moving on I want to spend a little bit more time in to audience targeting so there are two sub products within the audience targeting number one is we call it PG T which is also known as player group targeting and the number two is affinity it's also known as look-alike targeting so let's talk a little bit about what is a player group targeting what we do is that we look at the behavior of different are the behavior of the different devices in different gaming apps and group these devices into different gaming interest profiles and make those gaming interest profiles available as segments to advertisers or marketers so that marketers can have ability to run their advertising campaign to specific users with specific gaming interests that's what the play group targeting is the second one is affinity or collaborative filtering here what we want to provide is if we're given an advertising app we want the to identify the set of users who has deeper engagement in other gaming apps that are similar in nature to the advertisers app right so and make those users available for targeting with the marketer that's what we do in affinity so let's BB little bit into how those work so at a very high level we call this component persona it has four components to it it has a real-time screen it has a batch loader and it has the key value store we use DynamoDB for that and it has a serving side evaluation component we have we call it persona JA here let's double it a little bit more so in the dynamo DB or a key value store there are two tables that is a device table and a segment table the device table is indexed by device ID and it has a bunch of information for every app that the user has visited like the number of times that the user is playing that app the last time the user played the app the first time the user installed app the amount in the amount of enough purchases the user has made in that app right and the segment before we are talking about segment table let's talk a little bit about water segment is segment in a simplest form is a rule string that it captures a set of business logic that can be evaluated at runtime using the information available in the user table our device table that's what a segment is so for example for player group targeting we can define a segment saying that the it's saying that hey this segment will evaluate to true if that device or user has 100 plus boot ups in casino category meaning casino kind of apps and has some in-app purchases greater than $10 in again from casino apps so one leave the user has had a lot of game plays in casino apps and I have done some in-app purchases this segment will get evaluated to true for this user now if the marketer is targeting this segment then what that means is the marketer is targeting users who are big-time casino users that's how we use this segmentation and persona extreme to support this product a little bit more so before going into that let's talk a little bit about affinity targeting or look-alike targeting so here what we are saying is that the say there is an advertiser a who wants to target affinity users what he means what he wants to do is that he wants to find users who had a d7 retention and d7 retention means 7th day engagement so the advertiser wants to target users who are positive d7 retention in ads that are similar in nature to the advertising app so that's what the advertisers goal here is now how do we do that we first define a segment that is similar that looks something like this like there is an affinity score between - between up up - so we were that which determines how similar these apps look like my good limit for deeper into that and use that to find out if the user had an engagement in those similar apps let's find out more about the building blocks of the affinity targeting product so there are four blocks here there is a user profile again we talked a lot about that like this is stored in dynamo dB and we have a bunch of information stored by device ID and app ID in dynamo dB that's what we talked about in user profile then there is a targeting profile this is what the segment's mean here and I talked about how you can define different kinds of rules in the segment string right in the case of affinity segment we define a rule that targets similar looking apps right and that's what goes into a targeting profile and then there is an app to app affinity so we going to run a model to determine how we can group similar looking apps and I will talk a little bit more about that and then the final layer is device validation right a component inside an ad server that reads all of this and computes whether this particular device is eligible to show this advertising app or not let's double click more so what is this so we build a huge user to ask array right well here's one means that that particular user had a positive d7 retention for that app zero means that that particular user did not have a positive d7 retention for that app and yet means that the user did not install that app right so we build this array and from this array we compute conditional probability for a user to return on on day seven for an app a given that the user had a positive returns D sound retention on R P right so that's what we do here and the formula is over here a little bit more using this conditional probability we compute app to app affinity scores so here at one is cookie jam and after is juice jam and this under the affinity score of point six is the conditional probability of the user condition probably across a large set of users okay so this is how a user profile looks like inside dynamo we talked about that right it typically has all the information about good account last good update install date and a bunch of other information by device ID and app ID in the dynamo dB this is how a 30 segment looks like right so we typically want to say that this is the advertiser app and we want to capture similar apps that the app to the advertiser apps affinity score is greater than some threshold here in this case it's point three and use that to evaluate the users who had a d7 retention in those similar apps and then there is this Scala code that is sitting in inside ad server that takes the user profile that takes this target segment string and determines if this advertising campaign is eligible to be shown for this user that's pretty much the magic here that kind of summarizes most of the things that we wanted to say let me recap a little bit we started with saying providing a world view of chartboost and how advertising works a chartboost then we talked about the high-level architecture of chartboost the building blocks of attack and how apps are serving and data components are architected chartboost then we double-click little bit more into machine learning a chartboost an audience targeting product search ad boost which is where spark is heavily used and that brings an end to today's presentation of article chat boos I will open it up for Q&A you sign for me or is Dino for Magnus [Music] so let me repeat the question the question is that hey did you use a dynamo power segment that looks expensive so what happens is that the use dynamo to store the segments but we are not making any lookups as you said it is expensive so we read all the segments there is not going to be a lot of segments so we read all those segments once and get sucked into AK server memory and use that to process so we are not making a lookup to segment stable every time but we are making a lookup to use a table every time so the question is how our spark streaming failures are spikes kinda right so it's really again it's a depends on what failures are what spikes talking about but I thought a a couple of them that I am aware of so often some of the spikes happens we do to the increase and decrease in the data volume and we use the back pressure to make sure that the data gets queued up in Kafka if if we are having processing delays or if you are having spikes and data volume errors again it depends really on what errors we have Spears speaking about here there is a variety of errors that can happen right typically for most analytics pipeline what we have is that a streaming pipe followed by a batch pipe so that it's streaming run to do a problem then there is the hourly batch pipe that can catch up and fix those errors [Music] my dad you mentioned [Music] and everything the progression are using power relations through and ready to be [Music] okay so let me repeat the question the first question is that in terms of so simplicity we are using logistic regression and the first question is that do we believe that the model performance will be better if you end up using radiant boost or random for us that's the question one and the answer is that you do believe that gradient boost again we believe we haven't yet fully tried out we believe that gradient boost will help us out in the in terms of and will provide us some lift in model performance right it's early belief we haven't yet run tested it out and run in production yet also if you look at Kaggle papers and stuff gradient boost is one of the popular winning models in all the competitions right so we believe in that but we had again examples and competition is one thing running it at scale in a business is another thing and we haven't yet to reach that point then the second question was yeah what is the so you are talking about precision or AUC okay so you're looking for a UPR I believe we have I don't remember off the top of my head but I believe around 0.1 or something and I might be off yet with a specific number [Music] yeah I don't want to give you a wrong number I don't remember yeah yeah so again I mean this a good question I don't want to give you a wrong number I don't remember off the top of my head the training hall dimensions and he said the extract variable which might elaborate a bit more on the extractor variables of other men all my process and working miracles yeah so it is not a again it depends right so that is a so let me repeat the question so the question is that the feature is the feature extraction manual or automated right that's the question that's being asked again the answer is depends right so there's a whole set of feature extraction that we do that is automated from sparks collar code right and we use that inside our server because even the request that is at request coming in does not have the features that we need in a readily available form and we need to do a bunch of translation from add requests to the set of features we need and we use the same code again in our model training to extract those features from the data events and train the model right however that is like some feature translations or future variations that we do during a lab to find out if these features are going to be helpful or not right so initially we do that some of those manually but eventually everything is automated [Music] yeah we we typically do that with data scientists and data engineers acarbose go ahead your future and you're giving this up for you from now change what will change it's a abstract question I can go ahead three years from now yeah yeah yeah yeah so the industry the mobile gaming industry like in the next six months or one year we are doing a lot of work in ROI optimizations and LTV optimizations in one of the other university data science conferences I attended there are a bunch of folks from lift and other companies were all talking about how to do LTV optimization side LDV optimization is a pretty complex challenge that most of the folks in the mobile world are working on and I also see believe that more and more mobile and mobile gaming and VR will merge are well aligned and there will be that VR will bring in a new set of challenges that I can't predict what at this point of time so that's what I do any other question [Music] okay so the question is that water what is the key used in aerospike key value store and then how do we the second question was that how do we tag is the same user is using multiple devices right so the aerospike is also feed by device ID right IDF a that we get from our sdk and again aerospike stores a bunch of information for every device like the apps that the device is installed are the impressions and clicks and installs that the device has seen are clicked right so that's what is getting stored by device ID in aerospike and the second question you has was how do you tag this there are multiple devices for the same user so that's where cross device the mapping comes in today we do not do that like if the same user is having multiple devices you consider that as two different users our tool with two different device IDs right but there's a lot of innovation going on in that space right based on IP of those devices and a bunch of other metadata some companies are trying to map that these different device IDs belong to the same user or belongs to the same household but we don't do that [Music] I mean you when you asking granularity time granularity okay so question is what is the granularity at which we create models is that at an ad level or at a campaign level today we are creating at a campaign so we have about 30,000 campaigns but we don't create it for every campaign right this is where as part of simplification of the models we focus on the top campaigns that make the most impression and most revenue and group that let's stop them into a set of categories [Music] your English [Music] yeah so the question is that why did we choose a logistic regression the answer is that it is one of the simplest to machine learning models that while some complex models like grading booths or other other complex algorithms neural net random forests can cue a better performance bulk of the performance comes from using the model correctly using the right features reducing the noise in the model right and you can do all these things more easily if you have a simpler model and that was the reason for that [Music] two men worked [Music] so question is how do you troubleshoot and look at the data and graphs and do coefficient analysis when you have large number of features is that a question at correct so we look at it by each coil feature variable by different dimensions by different Geo's by different publisher sides by the different publisher group right so look at if you look at this entire room as a set of million features look at different bubbles and analyze that that's my simplest answers I mean we like to probably spend more time explaining what that means but look at different areas of feature set and see how the performances in that that bubble [Music] so teachers are speed [Music] - for machine learning so in the machine learning we are currently using mostly match that had trained every few hours in the audience targeting pipeline we are using streaming features [Music] and Jennifer's so the question is how many features are we using to train the model and not be using dimension reduction yeah so we are currently using million-plus features and we are using dimension reduction we are using frequency filtering they're using remove if we don't do anything it will be in the 30 million range right so we are using like we are doing first using something like look at top advertisers and top publishers and then groups are the rest of them into some slick category buckets we are also doing some frequency filtering right to reduce the number of feature set that that is getting loaded into a model [Music] it is the question is that on what basis we are reducing the features it's based on two things what other features are feature variables that are part of a significant number of impressions that contribute to significant revenue right and that I'm a lot of statistically significant data it's so wide the features that are sparse that are nice have grouped them into some bigger buckets that's what was primarily draw those feature detection I think the time when it's over so I want to wind up the Q&A session thank you folks [Applause] [Music]