Devreal

Building a Contacts Graph from activity...

Event: Scale by the Bay

Scale By The Bay 2018: Alexis Roos, Brad Powley, Building a Contacts Graph from activity data

Recording: Scale By The Bay 2018: Alexis Roos, Brad Powley, Building a Contacts Graph from activity data

you thank you very much Adam in the morning my name is Alexis ruse and director of machine learning are in sales cloud am here with Brad today to discuss our building a graph from activity data from communications there are two power intelligence services I will start by a brief says false introduction I would go over then white context models for what were doing for I and what we use a graph and Brad will then go over our platform explain some of the graph services we built and finally we will wrap up this presentation will cover some future developments and Salesforce being a public typically traded company I have to remind you to only make purchasing decisions based on products that are commercially available today that's for the disclaimer so Salesforce essentially built customer success platform that essentially allows our user a new to stripped all of their customers through a variety of interactions ranging from marketing acquiring customer cell-cell developing those customer service demo and so forth an asset in the last few years we've been back into adding AI right into a platform in all our applications to deliver the small the the world's smallest serum and we have made lots of investments and continue to make a lot of investment or an intelligent related company whether we're looking at bi machine learning or deep learning and the goal is to use the customer data we have and to make essentially or application smarter since force has a lot of customers we were over 200,000 customers and we have a lot of room for making applications matter whether we talking about leaf scurrying cache classification chat boards and so on and today we are specifically going to discuss what we do in a sales cloud reelected related to activity data so our team has developed a platform to enhance the CRM experience the customer relationship management experience using activity data such as emails phone calls tasks so we have some data in the CRM already but users can also opt in the activity data in our system and our platform automatically capture and federate those activities for user there also across users for given CRM opportunity so we have the ability to join all the activities for all the people involved in an opportunity or serum an opportunity will typically evolve like a manager or sales rep a bunch of faces and a bunch of customer people and we can still read that and we have developed email insights which allows to extract relevant insights in the emails in real time such as for instance is there the pricing discussion is there competition being mention is there and exactly being involved lowing essentially to to surface door in size so our users can take actions and can take the necessary you know appropriate actions based on those activities and we also support tracking of feedback and directly through some of the action we offer or directly as well to keep on improving the user experience today we will not cover but we will not discuss that we've done a number of talks on that section and for that part on the email inside side we're doing role NLP so these days we're doing a lot of work around like deep learning or an N Alice TM their walls but also using spark because we do scripture streaming for that and emails are complicated if they are recall on data prep so we have a lot of data prep to take apart the emails and be able to remove signature confidential notice rakitin before we go about scoring those emails and today we will talk about is like what we do at the bottom of the graph that we call contextual services which is taking those emails taking the communication data analysis we creating a graph to organize content knowledge about the communication data that we can in turn use for machine learning but also deliver new services and we'll discuss actually and Brad will discuss some of the services we built basically based on that data so I mentioned contextual services what is context and why does it matter for AI so in addition to be able to score those activities such as email in real time being able to use this tree caldera is very important as you can imagine it's important to score an email but also understand the context surrounding that email is it coming from you know an important customer prospect and so forth and all the leading AI apps all use context to improve you know the experience for instance Amazon will use you know contextual information about you to make you know their recommendations better and in the future in next generation I'd you become more and more clear that knowledge representation context we become more and more key to improve accuracy of AI over time so let's discuss all we organize and leverage some of the context very first so so we are dealing with you know enterprise data so there are notable differences between consumer and enterprise data and a lot of differences and a lot of things we have to do mostly in the enterprise the user isn't a product but it's a customer so as such we have all all capabilities that we have to support including there are security auditing but also retention privacy and GDP are in the trust is our number one value at Salesforce we host a lot of customer data we work with all customer Lena so we have to make sure that the customer can still retain control full control of their data in the impacts we acquired our manage data even creating a machine learning for those customers and the context will show in that case of what we're talking about here we have to stop it maybe at the user level at the team level or the organization level at best we just cannot share data across organizations because as you can imagine our customers don't want their data to be shared with other companies on you know competitors the context we have also is very rich in the case of the enterprise space we have like dozens or hundreds of different types of data that can be used to essentially know about our customers and in our case it's also very dynamic so we deal with communications there are so it's happening pretty fast it's changing in on being inhales you know dynamically so we have to be able to handle speed so for our customers in of the context enables us to deliver deeper insights this allows for instance to classify an email beyond just the content of that single email which is the oh well do I know the sender is the email essentially discussing my product my company product and services or is it discussing competitors war my organization can help sell to a given an evil or company can I get more background information about the company or the type of product and services they are selling identified the decision maker within the organization in automatically get an introduction so you can see that there are a lot of things that we can do once we have in a context about users and the organization now that you understand some of the benefits of acquiring leverage in context we will talk about what we use essentially a graph to to model that data so as you may realize that the graph is the ultimate there a structure to encode a complex relationship as I mentioned before we have a very rich set of data related to our user and graph is the natural data structure some of our organization may have you know hundreds of thousands of contacts or millions of contacts that extend beyond the organization because some of the Ovation will have in the next sets of ten thousand users and they'll have like essentially a lot of external contact and so they will have you know millions of events you know every week and we want to capture in or doors interaction over time and Moodle that as a graph in order to provide contextual services for machine learning but also derive new services and Brad will talk about some of the services we built such as for instance for a chimeric connection so if you look for instance at the top right we have a simple essentially image stream that shows you email being shared you know at various time you know between you know sender and recipient and you can see at the bottom that will construct a graph related to those emails and at the bottom we will have the CRM users the internal users and at the top we will have essentially two customers that are part of that communication in reality the graph is much larger right one or got at ham and will have mechanism to continuously update that graph and Brad will talk about that so the graph has we mentioned the goal for us is to generate context for machining one organ has more and more knowledge over time to know which even are which emails are relevant and we want to be able to use that to essentially reinforce all models but the email insights were getting based on the content of the email we can also push that back onto the graph so now we can enrich relationships not only we know people are connected but we know how they're connected right what kind of interactions you know do they have and and eventually are they discussing about pricing is it like engineering is it sells its marketing and so forth but then in addition to the graph can essentially be used in a standalone to deliver those new services so bright in particular we talked about recommend connection we organized that I know we can make ourselves rap you know smarter to why and go after new new leads or new people within an account and essentially by organizing contacts and by doing machine learning and kind of doing that feedback loop we can continuously improve the contacts and also the services we deliver in all ways on that context Brian will actually do the bulk of the presentation but before it a lot I just wanted to give one more thought on some of the work we are doing graph is great you know for modeling contains it's still an active active area of research as far our using you know email you know and deep learning on the graph these are some of the start we had so far for some of the insights we really worried about making sure those insights you know accurate and they also can be scaled and composed so one of the things we realized in early on was to develop context-free inside so what we have well we want to have contacts we want to develop contextual free inside so for instance doing an inside that would be essentially a pricing request might be difficult because a pricing request will involve no legend understanding of the product and services that the company is selling and typically if you're doing supervised learning and you wanna you know eventually label data is very hard to know what other product and services of the company so instead what we did is like we have an email inside that is essentially pricing description is the email about a pricing discussion and then we can mix and match that we've like essentially product detection because we can also look at what products the company is selling based on like the information that they have in the CRM and then and then we can allow composition as we mentioned we feedback the inside email insight into the graphs we had that feedback group and we can continually improve the way we do those additional services we've done a number of experimental Xindi planning graph a lot of folks right now doing things like encino edges or attributes we've done so more work actually on deep work and looking at random walk on the graph to do things like that person a person a similar to person be like see similar to person d so you can do some interesting you know associations there they are still it's still an active active area research there are a lot of challenges dealing with deep learning and machine on graph is difficult based on that the shape of the data but there are ways you can you can essentially create features you know create powerful features out of the graph thing camera over time and now we let Brad come over into the bulk of the the talk ok hello everyone so now we want to talk a little bit about the platform that we've built and also some of the products we've built off of it and so I'm really going to start with kind of the bird's eye view and we'll kind of zoom in to what products we've created and then zoom into some code around some simplified code around one particular product so try to give you the whole kind of notion of what we've done here before I get to that point though I want to acknowledge the team that's working with me on this product and platform that's NOAA Burbank and Gabriel Krupa these guys actually started working on it before I did so it is appropriate for me to recognize that fact that they've given this fertile ground that that I'm now planting in okay so here's the high-level view as Alexi mentioned we have people that opt in to Einstein activity capture if you do that and you're one of our customers then we crawl your inboxes and your calendars to generate an activity stream that activity stream in term intern gets placed into a Kafka queue and eventually persisted in an activity store which is basically an s3 bucket okay and so from there we get into this sort of peach rectangle which is where we begin to operate so from that activity store we take those meetings and emails and from that we build out this contact graph and you know we on board a new company from nothing just just its data but then as election mentioned you also need the capacity to be able to read a past graph from a different s3 bucket and then take new activity data and then merge it with the existing graph and then write out a new graph okay and so what you see here is that we are persisting graphs and we do work in in a batch process here okay so this is kind of the creation and maintenance of this raw graph but then also we need to deliver services and so from a raw graph we have a completely sort of separate set of operations where we're enriching that raw graph we're creating products from that raw graph and we're indexing data out to a store that our clients can access through a rest interface okay so there's a real separation of concerns here right we have the graph maintaining keeping it fresh and then we have products and services that we can build by enhancing that graph and that's sort of the theme that was a very conscious architectural decision we wanted sort of a fail-safe that we'd always have a raw graph we could go back to we also wanted the ability to branch off different products and different directions which maybe is is not efficient but it's it's very flexible ok so we what one I guess one other thing I want to mention about this is we just use kind of email and meeting metadata so we don't store the email bodies themselves in the graph for various reasons one of which is is size because it would increase the size of the graph by orders and orders of magnitude and so this is really about relationships between people but we don't have the email data and that's important it'll come up later ok so before we get into that I want to kind of zoom in to this updating and loading part of our architecture ok so again we have the activity store where the kind of the raw email and meeting data are stored and then we have the s3 the other s3 store where we have graphs stored and so from that store we have to take this activity data and turn it into vertices and edges in the graph and the vertices in the graph themselves are people right they're not just people but they are people they're people and then data about that person okay and so we consolidate those as tuples of an identifier a vertex ID and the contact itself okay and then the events are the relationships between people so each each event each email each meeting ends up being stored on an edge okay so it has an edge type what's stored within that edge is a collection of events okay so what matters this is a directed graph so if I send an email to Alexei then the edge goes from me to Alexei when he sends one to me another edge comes from him to me but if he sends five emails to me all five emails are stored on edge from him to me that makes sense okay so that's where we come up with new vertices and new edges and then we need to join that with the existing vertices and edges in the graph okay and we do that via we use graphics and so we use the RTD data structure for storing contact and edge information okay and then so from all that we get we go from this previous raw graph that's yellow to this updated raw graph that's that's freshly updated that's shown in green over there and we update these things roughly on a weekly basis we find that the graph data is pretty static it doesn't I shouldn't say it that way it's time constant is pretty long so updating on the order of a week is sufficient and it gives a lot of flexibility for our in 14 okay so as as Alexi mentioned each customer of Salesforce who opts in to Einstein activity capture they get their own graph right we can't have one giant graph I mean I guess I suppose we could with really complicated permissions for who could see what but what we do is we have basically a separate graph for each org that we service and because of that and because of the number of customers we're talking about scale does play a role so Salesforce itself has hundreds of thousands of customers around ten percent of which have opted in for Einstein activity capture to this point and for particularly large orgs we're talking on the scale of 10,000 users roughly okay so these are 10,000 people whose inboxes whose calendars are being crawled and then 1 million contacts and this is important because most people that we have on the graph we don't they're not our customers they're just people who happen to interact with our customers right and this will come up in OB it'll be important later when we talk about some of the features that we deliver and then finally a large or might have on the order of and I think Alexi mentioned this around a 1 million activity events per day many of which are not useful because their various forms of corporate spam or meeting rooms that are emailing people etc some of which can be interesting but many of which we need to filter away okay so our system basically has to work well in this sort of an environment and because of this scale we face kind of the classic trade-off of this tension between the size of this graph and what we can do with it right so gosh we'd like it to be as big as possible because then we could do lots and lots of things with it on one hand on the other hand the bigger it is the more expensive it is computationally storage wise engineering support wise etc gdpr wise okay so let's talk about one example of where this might be important and this is named resolution so when someone emails you you see their email address you oftentimes also see a name associated with that email right so you might see like if I email you you might see beep Ally at salesforce.com and then next to that it would say Brad Poway right so that name is what I'm talking about here now we might crawl a bunch of different names for an individual we often do so in this particular case we crawled various names for this fictional person his name is herb n' ulysses so we crawled that name we called you ulysses we crawled not odysseus and we crawled herbs okay so all these names and the question is well which one is his name what do you what do people think so for a person this is pretty easy to figure out right it's probably the top one and you could design a machine learned a machine learning algorithm to go through and figure that out but you could do this really accurately with just heuristics if you just had a little more data like what what kind of data would you want in order to be able to disambiguate which name it is what's that yeah is it like in a first-name lastname format you might also want to know for example did this did a particularly email where a particular name was crawled from did it come from that person so if I send an email to you well we're pretty sure if it's my business account we're pretty sure I got my name right right if you if if Cody for example sends an email to me he might have he might be using my nickname alright so that has a little less power yes right so there may be nicknames etc right so you you want to have you'd want to have a little more information in order to be able to determine a good name and you think well why is this important you know your users names right well actually like we talked about before there's an order of magnitude more people on the graph who we do not know who they are they just happen to interact with people who are our customers so we really need to get this right right so what kinds of information might you want you might want to know whether or not the name was using an email sent by this person we talked about that you might want to know how many events so we can have some frequency distribution over how often each one of these is used that might also give us a clue it's probably not as good as the first one but not everybody on the graph has sent an email right they may have just been receiving emails so we may have no data on the first bullet and you might want to know something about say the recency of use in the name and so this is just an example of data that may not seem important but you'd actually kind of like it to be able to keep it on the graph to be able to December you eight names okay so you do that and then you know from our system you get a resolved name and it's urban Ulysses and we're pretty sure it's right right so we picked a name based on that okay so let's talk about another example so this is in you know a lot of people are talking about privacy giving gdpr lately but we actually can think about privacy requests beyond just gdpr and here's an example so suppose you get a request your system at admin gets a request to remove graph data derived from events that include our CEO our board member and a competitor CEO and why might that request come in well maybe they're talking about an acquisition there may be some SEC reason you don't want just any user of the graph to be able to go in and infer that some big talks are going down right so so this request comes in what kinds of information would you need to know in order to be able to basically wipe these events from the graph what do people think you need to know which vertices are the CEO right you need to know which vertices are the competitors CEO also they ask you about the month of April you need to know when these events happen so that is these are data you want to keep because if you're going to be able to fulfill requests like these you need to keep that data so we talked about that you know you want the timestamp for the event you want some way some means of identifying all the people you want to be able to remove from the graph and so given that you can remove all events that include these CEOs and any board member within that time window and then if you recall edges are collections of events they contain collections of events and so any edge whose events were all removed should itself be removed so now all any sort of interaction between those people during that time frame have been removed from the graph okay so this trade-off is pretty straightforward when we're talking about something really concrete which is why I was asking you about this but the tension really becomes or the trade-off is much less clear for features you haven't made yet right so these are very very simple features two simple requests but you have to think about the future and like what might we be able to do with this and so that that's difficult you'd like to be conservative and to keep as much info on the graph as possible okay so the way we address this is by having multiple graphs so this raw graph that we've been talking about we just have that the one we persist we persist that with a large amount of data basically we're keeping as much of the metadata as we can in that graph and that gives us this optionality on new features so if it's there then we could then we can use it in the future it also allows us to branch more different kinds of features you know by building enhanced graphs off of that raw graph okay so we can like one kind of feature with one kind of enhanced graph we can branch another kind of feature with another kind of enhanced graph all the while never changing the route raw graph that exists and all the while being able to to failsafe to that raw graph should we need to okay so let's talk about that enhanced graph so an enhanced graph or graphs can deliver these insights Pacific reset results with faster processing so when we go to create a feature we take the raw graph we as quickly as possible remove irrelevant data from that from the raw graph that we need in order to make this feature so we can run very quickly and from that we can destroy distill this raw graph data into more useful summary information like the preferred name of the contact like event weighted edges so you know you just want to know the strength of connection between people that can be just encoded as an int right ultimately in many cases and also the contacts closest connections so these are just some examples of this so it's it's important to mention that we never persist any of these enhanced graphs right they just get created as part of a pipeline toward a product from the raw graph so let's talk a little bit about that pipeline so we've already seen how we get from s3 to a raw graph that's been updated appropriately what are the next steps well we do various tagging like shared account tagging so for example in many cases salespeople will use the same email account for for example marketing messages and things like this so there may be a single email address but you know a bunch of our names are on that email account like my name is on the account narcs name is on the account Cody's name is on the account Alexie's name is on the account and that's not a person the graph is about people right it's it's it's a set of people and so we want to be able to tag those shared accounts we want to be able to trim edges and lightweight the graph as much as possible that we talked about earlier so any information that's not relevant and material to making this ultimate product we want to get out as soon as possible in the pipeline we'll do some very basic entity resolution things like you know if if the like my email address you have one that's beep Ali there's one that's beep Ali one one that's beep Ali two etc we can do some kind of basic kind of email related entity resolution and from that we get this kind of partially enhanced lightweight graph okay so this gets us most of the way there but we're not to a product yet in fact we do some more tagging like a suspicious accounts so if you have if your name is like the name of something that has the word conference for a minute for example we want to tag that will do some even deeper entity resolution where we'll look at you know if there are two contacts one of we suspect they're both me what we do is we look at those two contacts closest connections if there's a strong enough overlap between those then we can get to better entity resolution and then finally we get to in org tagging so by that I mean in org is our customers inside of our customers org so it's not people they're selling to so out of work in this case is people they'd sell to we really want to make a strong distinction between people who are in org and people who are out of org because it's really important for the features that we generate and so from all of that we get an Enhanced graph we do product transformations on that enhance graph and then we populate the online index okay so we finally got through all of this now we can talk about some some products we've built on the graph some very specific ones okay so our first one is recommended connections so here's the problem we're back to urban he's a sales person who wants to get an introduction to Tasha this is Tasha and Tasha has an important lead he knows that Tasha's someone he wants to get an introduction to and how can you go about doing that well that's where our recommended connections product comes into play we can look at kind of the strength of edges between Tasha and her nearest neighbors and that strength can be related to timing like how recently they've interacted whether or not they've had meetings versus emails and weather that they've have reciprocated emails so in some cases you have sort of these windbags where they just kind of keep slamming someone and they never get an email back that's not an indication of a very strong connection it's it's strong as soon as you get back it's like this is an indication of a stronger connection and so from that we can give kind of some scores to each of Tasha's closes connections to see which ones are the best connected to her okay and from that we see that Andrea is the closest connection to Tasha and so the solution is we serve to urban that Tasha's closest connections are Andrea is the strongest and Baskar is the medium connection he knows Andrea now he knows who to go to to get the introduction to Tasha that makes sense all right so let's talk about the oh sorry one other thing we sort of have a value on our team that wherever possible we try to serve a rationale as well so that it's not so black boxy right so we can give a rationale it says Tasha and Andrea emailed each other several times in the past month like why do they have a strong connection well this is why this is why we say that most likely not right yeah so Tasha might be someone that urban wants to sell to mm-hmm yeah powerful distinction I'm sorry the question privatized say more about that yeah so there may be and that sort of gets back to the feature we talked about earlier about you know let the CEO example and do you want to be able to you know know that that like for example if I can click on Marc Benioff and I see oh he's emailed and what so there has to be a mechanism to be able to remove that information from the ground yeah let me let me table that until the end just to make sure we have enough time if that's okay yeah thanks all right we're close okay so the next two out of the three products will mention this one's called contact mentioned and this is sort of a mix between email insights that Alexei talked about and contextual services and so here's the problem so urban gets an email from a client that mentions you go and you go is a person that that urban does not know well okay so here's the email and it's body and again this is part of our structured streaming email insights side of the house so the email says before you send a quote to Sarah we should first convince Google that your product can help service coach Eve its goals okay so through named entity recognition we can figure out that there's a name here and the question is what do we do about that name is we have someone that's important is it someone that should be mentioned and so the solution is that we do this kind of automated review of each of the participants connections okay so we look the participants are urban Sarah and Raia Raia was the sender we look at Ray as closest connections and oh it looks like Google could be Google Garcia right he shows up there we look at Sarah's he shows up again and he doesn't show up in Urban's and so we can say we're pretty sure that you go in this case is you Garcia looks like you've never interacted with ago so we can say to you would you like to view Google Garcia's contact information it was a quick like one-click method to get from this email to the person they're talking about okay so that's contact mention and then the third example is influencer or identification answer the problem is and we talked about this before so urban wants to get an introduction to someone at widget corp now in this case he doesn't before he knew he wanted to talk to to Tasha right but this is the use case where he knows he wants to get in an organization but he doesn't really know where to start right and so he wants to get an intro here this company has high potential but he doesn't really know where to start and so the solution we have is kind of a two-part solution part one of which is that our algorithm our influencer algorithm finds a person who is both well connected and is showing a willingness to reciprocate was in the graph and in doing that it identifies that she for node X here is the influencer at widget Corp and it gives a rationale so among a group of five people she is particularly well connected having sent emails and met with three people okay and so that's step one and then step two is okay so this is the influencer so the logical next step is is well how do I get an introduction to this person right toom do I turn for this intro and so now we go back to recommended connections and urban can learn that s Condor his coworker is well connected to she and so now he knows like a set of next steps to get to she can I table that just until the very end yeah thank you okay so let's see how much time to have five minutes all right I'll talk about this very quickly so we have three methods to identify influencers the first one is we needed some way of separating this large like universe into some small worlds in the graph and so we need something that will enable us to make cliques in the graph we also then once we have those clicks we need to figure out like how to score someone's influence or nests right as you might call it like how strong an influence or are they in that click and then once you have like a distribution of scores over the click how do you go about choosing which person should be deemed as an influencer so we have this namer function as well okay so for this click maker function this is I'll mention this because it's interesting so our graphs tend to be these scale-invariant power law graphs and researchers that do work in these know that the way to break these sort of networks up is to find the super nodes and remove them okay so this is the way to do it people that look at you know breaking up terrorist networks etc use these kinds of approaches well it turns out that who are those people in our graphs they're our users right the people whose inboxes we're crawling are the ones that are the super nodes in our graph so we already have tagged which people and that's why this in org tagging is so important we've tagged which people we need to remove we do that we can get to this graph full of cliques and small worlds and so I'm gonna skip through the score and the name er I think that's pretty straightforward the score puts scores on the graph based on the connectivity in this case I have just like a one-liner using page a reverse PageRank to be able to figure that out you can use any number of algorithms to show this and then you need some means of saying hey you know above some threshold we're gonna call this person an influencer okay and then of course we want to give a rationale at the end why is it that you've tagged this person as such okay so let's wrap it so one of the messages key messages from our talk is that context gives us a deeper meaning to data it's good to have this context we can do more interesting more meaningful features for people who use our products when we understand kind of how they interact the next is that graph is a powerful data representation and it's a great way of encoding context and effort some complaints about spark here this week we're you know nothing's perfect but we feel like Apache sparking graphics I've done a reasonable job for us and it provided a reasonable platform for us and providing these contextual services and finally I think what Alexi talked about this sort of synergy between email insights and contextual services really go well together so we don't want to save email bodies on to the graph but we could save insights about email body so we could say like hey this edge contains a pricing discussion right and then that helps you give better products based on contextual services it also gives you a prior distribution when you're trying to score an email so you might say like hey we're kind of on the edge as to whether or not this is a pricing email but we look and we see that there have been many pricing emails that have gone back and forth between these two people and so that helps with that so that we're about at a time and maybe there's time for a couple questions ok great Joanna do we have a hill [Music] particular customer because your customers are the salespeople and they write to many many people yeah that's exactly right and and it's just like that and we have very little information about those people we only have information based on email metadata and what people talked about so I'm sorry I maybe I didn't answer your question can it was it a question in there that I missed know how you can help them okay right so so presumably removing the spam the spam that you removed presumably is not relevant to that to the influencer relationships so for example whether or not you know like an automated email generator is part of that customer organization doesn't help us understand who the influencer in that organization is is that answering your question I feel like I'm not I'm licking your face and I'm not seeing clarity what's up like how we would write yeah yeah yeah yeah yeah so kind of a cease-and-desist email that's what you're saying so yeah so I think the way that we would handle that is through email insights so there might be you know NLP we do on the message that finds these kinds of cease-and-desist emails extracts there that summary of them and then we add that to the edge on the graph based on that so just like I talked about you might have like an insight that that reason email and says hey pricing was discussed we'll put a pricing tag on one of the edges in the graph you could do the same thing with like a cease-and-desist email in case we this is the one too many wealthy people are involved versus meetings with Mexico City don't know if he managed recently so they are all black we can import to kind of like put that kind of interaction underneath but you can have false positives like with any you know intelligence system that's right there was a question in the back someone's been head to head retention and I think that's sort of right right right I don't know do you want to take that huh okay okay I'm not sure I got the first question hundred percent but from a 16-point will coin goes in every time and they are two fools to throw at the top is processing those imaginary time to deliver those insights because we serve that who set about but also foreign concepts washing box so if your sense of sin box as you get sure it many you get to know immediate identification on the email that for instance is as kidding requests that you can take an action so we do process that in real time 14 early insights for the graph we do it in boxing the reason why is because normally we build a graph but we do those higher different calculation on top of the graph that page rank in on a recommendation in the expensive you typically cannot do those a person in real time right typically hard to use any kind of large-scale distributed database and do a graphical permissions in real time so for that we do that in much and then we index it so can be searched online asking do we have time for any more questions or said okay so the same revenue I save one more time [Applause]