Devreal

Getting Healthy with Spark: Big Data at MyFitnessPal

Event: Getting Healthy with Spark

sfspark.org: Chul Lee, Hesam Salehian, Getting Healthy with Spark and MyFitnessPal

Recording: sfspark.org: Chul Lee, Hesam Salehian, Getting Healthy with Spark and MyFitnessPal

hi everyone so my name is Julie I tried to speak up but if I can then my apologies so today I'm going to talk about how we are using spark at my face both together with Sam and then I'm going to start with some kind of cool intro of our company and then expression in the context of under armour because we go quired this year so after i'm done with that i'm going to talk about some data assets that we have and then after that we will talk about spark at my face pal why and then some quick summary of some use cases that we have and after that so my part is actually going to be some how boring it his part is going to be more exciting bit because it's going to do some kind of deep dive into details of how spark is being used at my few spa and then after that we're going to have some kind of trade sessions so feel free to ask questions so okay so what is what is my feasible so how many of you guys have her about my pins power okay so pretty pretty number that's good so we are simple and effective health and fitness tracking tool we are one of the one of the best known health and fitness apps and then we have which I mean this huge reach of you know 70s more than 70 countries so we have we serve like pretty international kind of user base so and then and our users having really like you know loving a rap right so because of this we have a very massive of and very like suddenly like committee of users which is more than 90 million users and this number is going right one thing that we are very proud of is so we have one of the largest food database and then this is very interesting kind of thing in the sense that we have been accumulating this big database or knowledge base of food items and first first time in the history of mankind I like and then this is really amazing right so no one was able to really like you know archive or track this database of food items and then my friends pal actually did that and then are luckily enough because of this we have been able to accumulate over 7 million food items right that tends that translate into 19 billion for entries so this is a kind of very interesting you know database of really like being able to track how how users and how people are really consuming foods and then like what kind of diet patterns they have and so on that's very interesting right so as I said before my face pal join under armour in march of this year and then after that oxidation we became more interesting because now thanks to Under Armour like our reach has button right and then now we're talking about 150 million registered users cause like you a connectedness because under armour about like got two more apps before my fitness pal and then together with my fins file my fins fell user base and then the rest of other apps like we're talking about a hundred fifty million users right so so this is like article that talks about all like how under armour is trying to become a tech company right as you might know under armour is really like emerging brand will it going happily after like a nike right and then you know getting lower attraction right so but at the same time an underarm was trying to become tech company and then this is why i like you know under them about like different apps and then that's why i like we became part of under my family so what does it mean so this means like we have lot of data right in various formats right so that means like because of my face while we are you know pre in reaching chumps of food and nutrition data right a lot of food entries lot of food items i said before and then you know 38 million recipes right which has been generated by it on 190 million my original users as i said before right but but at the same time let's say before because of other apps we have a lot of workout data right so now I meaning like running data where we have a lot of time series data we now have music data with work out right we had you know walking data we have activity data we have sleep data and then we have a lot of data that will be able to collect you know to our partnerships and so on so we have a lot of data right a workout data in particular and on top of that now under armour has been a retail and eat company for for many many years before we came aboard so therefore like under armour is actually had a lot of product transactional data and then in addition to those transactions data or we have a lot of user preferences on different clothes shoes we're with you Isis and so on right so you can really see the the wide spectrum of different data sets that we have right so this is very exciting so why spark right there's there are several reasons why we started like consuming spark as an option for our data warehousing so number one was like you know data volume right we have a lot of data right so okay so we had to think about ways of processing that volume of data right and the same time we had to deal with all kind of different different kind of you know data as i said before right so we had to find a tool for that right so then in that case we pretty much have two options right you and then more traditional way of doing that would be do right that's what everybody has been doing right the other option that we had was like gonna spark as a matter of fact like we ended up really using both so we use both in hydro and spark right but the reason why we started like looking into more a part was are all changes were a little bit different from in a traditional sense of or like map use type of thing and the user store the data in your head of you know cluster and then you do typical each of stuff and then on top of that you do some kind of error processing why because we had to deal with a lot of you know complex and diverse like argument needs right more specifically we had to deal with i got a memory state aspect of data meaning like when we are doing a lot of you know data processing and then machine learning stuff or like development of data products we had to persist data memory right and then as you might know already when you're using hadoop right it works when you're like using hadoop as a simple storage engine and you sort just data but then you want to use Hadoop as a computation engine it kind of works but not it doesn't work all the time I especially when you're trying to like hold your ears you're like a memory state right so one example of that could be well when you try to run lda stuff right sure like using Hadoop you can do that but each time that you're running like your LD a computation then you have to really like deal with a lot of disk i/o like botnet because like you can persist data there for like you have to store that persistence in in disk so probably does not idea right and they're like numerous these cases for like you know why like in memory still kind of machine learning or like your data processing are desirable right at the same time like we had this kind of speed requirement right so because we're in the disk domain of health and fitness right so it's not true all the time but most of times faster they're processing we result into a better customer satisfaction and products perience right so for instance like if you want to like provide nutrition and running inside ideally you want to like react as quickly as possible where and I just wait like a week and then after week whether the value of your inside would be really you is going to very diminished right so we want to like I have faster process in there right another example which is more in the context of retail would be if you want to run some marketing campaigns right and then obviously you want to be very reactive right the moment that somebody is purchasing something naw moment we want to really like target and that user and then you know provide some kind of marketing campaigns right depending on like some transaction data that user has provided right so that's just as well like you know we start linking into a spark way I'm not saying here that sprach is the only way of really solving all these kind of problems and I'm not saying that like you know how dope is going to go away right so but I'm saying is like for our like on purpose like since part has to be very effective right this is why we are here to share some of our expenses okay so as I said before like we we are using spark get my fins pal and then we are expanding our use cases the first use case that we had a vs part wells machine learning right so this is quite interesting because if you again like think about the evolution about how to infrastructure right so the evolution how do in France mean okay data warehousing first so you store data first and then like you build a competition engine on top both of a dog right but but from my understanding like one one thing that was very interesting about spark was like spark was born as a computation engine right and we I know we have some some spark guys here so correct me I'm wrong but that was like you know kind of the emergence of the birth of spark so you started as a competition engine so we have this like luxury of okay really by passing this data warehousing aspect of Hadoop and then like directly start doing a computational aspect of data right and that's what I get the first application that we had for sparklers machinery and I and we felt that was kind of natural feat for a spark and that worked out very fantastically for us and then that's when we use a ml deep as up like a way of doing machine learning and then in that process be a lot of help from data breaks and then saw inventors of him a leap right the the other use case of spark and my principal has been advanced data processing again I'm saying like I'm I'm putting like advanced like in in course because you know like using you know Dube or any kind of traditional data warehousing or any kind of databases you can do data processing right so so I don't think it's true that some people I don't agree with some people claim to say oh like spark is going to really like override like Hadoop and they is going to place Hadoop blah blah i don't i don't believe that because I mean Hadoop has his own place right because for simple data warehousing definitely like how to be useful but when you want to start doing advanced data processing right I think sprog is definitely efficient because you can store like again you're steady memory right and then you can persist right and then that means like once that once that you want to do our advanced data cleansing then persisting that state memory is extremely helpful and an example of that could be if you want to do some kind of clustering stuff then you want to maintain some your your custards new memory and then then for that kind of proposal and doing a memory they're processing is definitely help right advanced data expression analysis same thing like yeah you can just claim that all like you can do aggregations and you can really store this aggregate is like using Hadoop true but if you can persist those states via memory write faster more efficient and then more scalable right some Attila up saying I like you can you can like your ETA jobs using loop right but if you want to do like a lot of batch processing with some states are kind of persistence then probably like a sprog is the right way to go and then that's kind of the area that we exploring right finally we haven't using like spark for product and decision science right again I don't think this is really mandatory because you can just same with a hard infrastructure or other like you know data infrastructure but sometimes you need like speed it right and then sometimes you need like a memory persistence again by then for those kind of use cases definitely you want to really have this tools handy so that you can really like have a very quick like turn on like you know one like probably incision sciences right so this is the way that we haven't used spar get my visa and I think some is going to talk more about all right hello everybody my name is hossam i'm a senior data scientist in my fitness pal today i'm going to talk about a sample use case of using spark for our data processing which is verified food it's a common feature among our users because we've had these requests from users for a long time to come up with the most accurate subset of the foods that we've had in our database with respect to the nutrition information and the reason for that is be based on the crowd-sourced data that we've collected from users we have all sorts of duplications and inconsistencies among our dataset frames for instance we have like mcdonalds mac chicken sandwich in lots of different formats spare especially if you consider different misspellings like the first example there are several items referring to the same entity in our database and from the nutrition data it's kind of inevitable to have inconsistencies event for the items that are referring to exactly the same thing for example these are the two sets of nutrition facts for apple from two different resources but we can see that the nutrition information are different from each other so this is very important in s our users search experience because once we present the search results you can see that for example for the item like tangerine user can find all different source of calories and like it might bring a lot of confusion to the users what item to select and uh more importantly they may not have much more confidence to the existing nutritional facts so that's that was the number one motivation for us to go after this really accurate subset the other one was the existence of a lot of private foods in our database by private I mean the food items that users can just create themselves and use themselves and these items are not necessarily searchable so based on what I described earlier we have around seven million food items which are public and searchable but we have a lot more food items that are private and before doing this project we hadn't used this huge set of food with for any purposes so that was a very good opportunity for us to start looking into these private items and having an aggregation process to compute the most accurate nutrition information so the goal that we were trying to target was to have this kind of search experience such that you can see this green check mark for the items that who's nutrition information has been verified by this process and this is very ideal search experience that we were going after and the solution that we presented was based on a process in a spark where we decided to use implicit and explicit signals from users so there are several types of inputs from users that we could employ in order to compute this accurate set the implicit signals are the number of times a particular food item has been locked and that would probably tell us something about the correctness of the nutrition information and also the convenience of the serving sizes for the food item and also if users have created a lot of private foods with exactly the same name the same nutritional information so that would tell us that they probably are referring to the same resource and that's kind of like a trust of a resource for us to make the nonsense based on and the explicit signals are some of the signals that we've gathered over time especially for example if the users have confidence on institutional information we can get some confirmations and we can consider the number of confirmations for food item as a very strong signal but unfortunately we don't have many of these types of signals is kind of sparse and another thing is the public and private foods which are kind of treated in a different manner so these are all the signals that we were dealing with to come up with this process so any question of the entire attrition I'm gonna call like very scared taxes black white high contrast and when I was just one ah ok so the question was since we already have the feature which is barcode scan why don't we have the feature for the users to take a picture of the nutrition facts and just do an OCR on top of it to get the nutrition facts but the answer to this question is the most problematic examples are like apple or McDonald's Burger the things that they don't have neither barcode or the nutrition facts on top so for those kind of stuff we have a huge set of imported food items that's why we compared the barcode with the one that users scan and we make this comparison so we already have the nutrition information for the ones that are that have a bar code attached so we don't need to like do an OCR on top of them to just get the nutrition information but that was a part of like a subset of the all the verified foods that we have released so the subset that I'm going to talk about was the ones that we came up with by processing the food in our database so the as mentioned earlier there are several benefits of using a spark for our data processing we had a different source of data from different types and having them all together and doing some kind of like joints on top what's very convenient in a spark and also there are some libraries in a spark being implemented like ml lip and graphics and that was very good for us and that is very convenient for us to use and also a lot of process that we're parallel in nature could be executed in a very fast manner so in the next few slides I'm going to walk through the whole process of finding this subset of verified food the data that you are dealing with us as subs of a combination of all source of data from nutrition information from the foods the string of brand description the user IDs like who's created the foods and also the signals that I was talking about earlier which was the number of times each food has been logged and etc so in spark we could just simply register all these tables together and we could make a join in a very fast manner and by the end of the day we had a class defined called food and it had all different fills that we're looking for for example this string of brand in the description the nutrition facts like calories fat 14 etc and also the signals that were coming from users so the rdd that we were dealing with afterwards was just a simple rdd of the foods so that was the natural way of dealing with this huge data in a spark so in order to make the process more accurate and since we were dealing with the user input data we had to run some kind of string processing on top and this is drunk processing is very simple string manipulation like removing the punctuation characters or duplicate words and also sorting the words in the Brandon description based on the first character because in several cases the food names are referring to the same thing but words are kind of different order so on top of this food rdd we had to run this normalized brand plus description and what we'll end up with after this step is an RDD of key values which are a string which is the normalized Brandon description plus the food item itself this was very interesting for us because after this we could apply the simple function implemented in skala which is grouped by key and to find all the items that are having pretty much the same name that was probably the most important except that we could take care of for our aggregation process so there are some instances in this picture for example coffee mocha starbucks is a collection of all foods that are actually referring to the same entity and the same example for like McDonald's nugget and etc and if the names are kind of weird is because the words are kind of sorted so they may not make sense so after this cluster identification we had to find one of these members of each class that is the best representative so the definition of best representative is really tricky especially when we're dealing with user crowd-sourced data but based on the signals that we that I talked about earlier like the number of times each of these members have been logged or have been confirmed by the users and whether or not this food is public or private we decided to pick the one that has the highest score and this score is computed considering all different signals so that's why this particular item is colored in green and from now on we are trying to get the nutrition information from this cluster representative and to make the nutrition information as complete as possible based on the rest of the items in the cluster the one thing that I want to emphasize is that these clusters may be a good enough meaning that we have enough number of members in each class and we have some level of confidence to the nutrition information that have been aggregated from each cluster but in some cases this classroom is not incorporate into generating a verified food so for the items that for the classes that we have a verified food so for instance this green role was the item that was picked as representative but having the nutrition information even from this representative it may not be tras Sybil and we decided to go through the rest of the members in this class and fill up the gaps for example if the fat value is missing in this representative based on some kind of voting and aggregation over the different members of this class we decided to fill the gaps and return the most accurate and most complete set of nutrition information so one step that was kind of mandatory for us to take care of was to get rid of the duplicates so after running this process and after computing the verify food for example we realize that we have four popular items we have different versions of the verified foods and they're all promoted from different clusters for example for McDonald's chicken nugget we have different types of the strings and not all of them are necessarily and end up in the same cluster based on the string processing that we presented so one important step to take care of these duplicates is to run a dee doo process and we decided to run this to due process on top of the clustering asset and that's why i have put at the step 1.5 because it's between the clustering and a nutrition I gregation process so in order to run it we decided to do it on top of the clustering for example if we have a cluster for coffee mocha starbucks and we have another population which is probably a smaller population but they have coffee misspelled so they're referring to the same entity but with the string similarity it's it's not really easy to find these instances so we decided to have this deduplication process and once we figure out there actually referring to the same entity we decided to compute the union of these clusters and generate a bigger class and we doing so we can benefit from two advantages at the same time first we can get rid of all these misspells and the words that are actually duplicate and also we will have larger classes meaning that we will have more members in each class and more members provides more confidence towards the nutrition aggregation process so that's an important except that we had to take care of and in order to solve this we went after a pair of eyes a string similarity comparison between all the cluster representatives but as you can guess this is not the most efficient way and that was quadratic with respect to the input size the running time was a lot for 1 million input foods the running time was kind of reasonable like five minutes it took around two hours 45 million items still not too bad but when we wanted to try for 50 million it could raise to more than a week we actually we didn't try to run it we just estimated that it's going to take for a week at least and we're not that patient to have this running term I'll do this process was extremely simple and the main downside with this pervez comparison was that it's not paralyzed about because we cannot like separate out different computation parts because all everything is so relevant to each other in order to overcome this issue we made use of locality sensitive hashing which is a well-known technique in machine learning community it's abbreviated as lsh so lsh is a method that in a probabilistic manner predicts the members that are with high confidence their belonging to the same bucket or the same cluster for instance in this case given all these clothes surnames lsh could help us identify the candidates of being duplicated and of course since the process is not deterministic we will have a lot of false positives and false negatives but considering this speed of this process v that was a huge game for us and the fact that the whole population is divided into different buckets was very appealing especially when dealing with the spark because we could paralyze the process for each bucket and run a string pervez the string similarity within each bucket only and this process was a lot faster as you guys can get and it took less than five hours for the entire data set so considering the one week versus five hours that was a huge time game for us and actually we use the elastic coating swag which is available in the github and that was very efficient for us so on top of the nutrition aggregation and extracting the verified food we had to add another layer as a post processing step the reason was we ended up with some verified foods where you can see that they only have the calories information and the reason for that is most of the users don't really care about the macronutrients or micro nutrients for example fats carbs sodium etc and they only add the calorie for the food that are trying to log so we ended up with several examples like that and since there is a relation between the macronutrients and the calories we decided to add another post processing layer on top of the verified fruits and we applied some very simple rules on top not one of them was having non-negative nutrients for example we realized that several verified foods were called breastfeeding evade minus 500 calories and I'm not sure if it's correct or not it's accurate but from the users experience it's not very appealing to have this item and have a verified mark on it and another sign of the check was the relation between all the and fats with respect to the total fat and the summation of all should be less than or equal to the total fat and the same for the carbs sugar and fiber and the last one which is a little tricky is the relation between different sources of calorie with respect to the total calorie with meaning that the total calorie should be pretty close to the weighted summation of carbs protein and fat and here is a snapshot of this last sanity check on top of the verified foods we can see that the calories protein carbs and fat have been retrieved from each of these verified food and this relation has been used for filtering the ones that are I mean filtering out the ones that are not satisfying this rule and you can see that one the ten percent margin has been added for this equality constraint and that was it so after by the end of the day since most of our users are pretty obsessed with before after pictures i'm adding this before after picture before having this feature and afterwards and you can see that only the items that are referring to tangerine and we have the most confidence up on are marked as green a verified item and the rest of them are pretty much the ones that are that we don't have much confidence so users are going to much easier if selective items that you're looking for and that was it that was a very interesting application of spark in our database and if you have any question will be happy to answer right we also stick so for 16 long starbucks coffee right if it says it's coffee advances we going what happened to see oh okay oh you mean four different serving sizes sure if I understood the question right the question was if we consider different serving sizes for the items that we were trying to aggregate for the nutrition information was is it correct so in that case since we only consider the ones that are having the most agreed-upon challenges information the chances are the serving sizes for these items are pretty much the same but we actually decided to pick the serving size that is most dominant among each cluster so between the items that are having the same color information we decided to vote on the serving size that each item has so we had this kind of aggregation process for both calories and serving sizes so by the end of the day we are pretty sure that this item with the serving size has this particular calorie information it wasn't like a speed issue for us the reason that we actually did this simple group by key for clustering was that we wanted to only put the items in the same class that we're one hundred percent confident because for any clustering algorithm there are a lot of inconsistency in the abilities and we are not hundred percent sure that they're not going to be false positives or false negatives but in this particular example we wanted to just reduce all the items that are having pretty much the same name so that's why we use that and we employed this duplication process on top but one another alternative was for example to use the clustering or k-means clustering for instance in ml lib that was another alternative but we try to be more conservative oh okay the reason that we used that was we actually implemented sorry the question was was there any specific reason for using our DD vs data frame we did this project like in january-february last year and we didn't have like the data frame so that's one alternative for us to migrate this that's good question so yeah the question was how many data sources we had what was the type of the data that we're dealing with so we had some data living in s3 one was for example the park a five was in all it was one format we had see CSV files coming in we had some data in TX file so and also as far as I remember that we had data in our seat tables and like aggregating them all into the spark was very convenient because we had like the tool for Linda not for this project but V for other use cases we have like five tables that we need to aggregate this therapy a little gate you I guess that's actually the goal it's the Indians current form it's not but we have the closer information being a sword and we have the information and the signals coming from users that what verify foods has been logged more compared to the others and with respect to these items we can just make some improvement to the algorithm with the user information it hasn't been implemented that what we have the infrastructure to employ that keep it for or the question was if we sort the data before duplicate deduplication or we just didn't so we actually didn't store it we had the whole data after clustering in an RDD and once we deduplicated the data the algorithm for doing that was to just merge this cluster so we didn't lose any information we just reassign the cluster for example for the guys from this misspelled version to a new cluster representative so we didn't lose any data we just change the cluster oh the question was if we measure the accuracy of clustering the answer is since we didn't use like machine learning based clustering in this case it was like deterministic but if you mean the lsh for the lsh the accuracy was around seventy-five percent or so meaning that seventy-five percent of the times that like a bucket was detected by lsh it was actually a to duplicate but for the general clustering algorithm since we only use this group by key operation it's kind of deterministic there is nothing to be worried about it depends on this string normalization for sure so the question was if we had some evaluation on top of this process so are you asking about the correctness of the nutrition information of the data or you move just other metrics computer Oh that's a very good question that's actually we haven't looked at the matrix yet but we would definitely expect the users to I mean for example time prologue or search padlock to decrease significantly but we haven't measured the stats yet but that's a very good question can you repeat the question what do you mean by normalization please I'm the question was how do we normalize the user signals for verified and non verified food so it's kind of a follow for the earlier question since we haven't looked at the metrics yet that's a valid concern because users might be more bias towards verified foods versus non verify it's like the first rank results in the general search engines so in that case we haven't thought this through but that's one of our concerns but in general we decided to not to rank I mean to I mean rank the verified foods too much higher compared to the non verified items because we didn't want to mess up with the users I mean usual experience so that was like another concern what part sorry okay sure that was a good question the question was a little bit more details into the LSA algorithm if you go through this github link you can realize that this illustration algorithm in the spark has been implemented for integer values and the main concern for us was how to make this transformation from string features to the integer features and actually the features were like sparse vectors of integers so in that case we created three grams for the characters of the existing in the strings for example if we have coffee mocha starbucks we consider all different three grams three lengths strings appearing and having a universal vector we specified them I mean we can't be counted up and down the number of times it has appeared so we ended up with sparse vectors of integers so that was like a natural input format for the other set and the output was given this index what was what is the pocket index that we that is fine from the lsh algorithm the vector was did the entire set of possible three grams for example is starting from AAA to Z Z Z and since that's a huge vector of size I think seven seventeen thousand or so only a few of them are nonzero so that's a sparse vector and that was exactly what this message code was looking for spark sure why ah the question was using this transformation if we bias the longer phrases too much the answer is if we have a longer phrase and there is another longer phrase that is pretty similar to that so that's natural to end up with in the same cluster because we have a pairwise comparison by the end of the day and if we have a large phrase and we have another small phrase they're definitely not referring to the same entity and it's appealing not to end up with the same bucket that's true I mean for any Austrian similarity I mean the even with the pairwise comparison we were using a normalized edit distance and it's normalized is divided actually by the length of the string so for larger string we have this problem for sure but I mean sometimes it's inevitable to face these issues this part very yeah actually it was taken care of by this normalization in the sort by words so based on the first letter of each word they have been sorted for example you can see that here rather than having subway meatball sandwich footlong we have the whole string being sorted and that's why it's not making so much sense it's like footlong meatball sandwich subway they're afraid the same exactly they're going to end up with the same class here and the reason was for the crowd-sourced data we had cases that the user put the description into the brand and brand description and just having reverse orders for the words industry so we had to deal with all of them and I mean that was like the earliest a force that we made in order to make them referring to the same entity oh so can you okay so in those cases are the divorce case would be one of these entities with one of these words they're not going to present a verified food I mean the verify food would be corresponding to only one of them which is like but at by the end of the day since we pick the class representative the right order with the right nutrition information would be selected exactly yeah so in that case yeah as I said the concern was if we have two different entities that are ending up in the same cluster so in that case the worst case scenario would be there won't be no any verified food from this cluster or if there is only one of these foods are going to have a verified food version and the other one is not going to show up because when we have this voting scheme in place I mean that vote is not going to be counter because it's like the calorie information is like different so the question was if we have translations to other countries we act didn't explicitly translate foods but the same framework the same algorithm was independent of the fact that the words are English so that was the case for any non-english characters so we ended up with a population of non English food items that have been verified since we have lots of non-english foods in our database like there were like enough number of them sometimes that could promote to a verified food so we had some items but we didn't have an explicit translation on top of each verified free kind of a rush any kind of like perfect accuracy use the navigation let's say if you had sufficiently accurate data or some of the research question given with food particularly to call macro-evolution reporting on their body fats and it waits with the what data and help Trisha Paul research what particular target resulting doing loss could be so that was a very good question and it was too long for me to repeat I I try to repeat some part of it so the question was if they assume that for example food information that we have in our database and the food log events coming in from users what are the post potential flows to do researchers for the users for weight loss or weight gain or food habits and this kind of stuff so that's for as a data scientist these are the really exciting projects that I can work on and especially as tool mentioned after being acquired by under armour we are now enjoying with MapMyFitness app which is like a map to keep track of exercises by the users so now we have two very general sources of data the nutrition and exercises and of course there are a lot of relations between them so that's something that's actually the main focus of the entire team working on and as you said that's based on the strong assumption that the data that is coming in is not noisy which is not a good assumption in general but there are a lot of classification a lot of recommendation etc has been going on for instance in the in this year's spark summit in San Francisco another teammate in my fitness pal presented a method for recommendation using spark for our food database based on the user's food consumption history so there are all sorts of these data analysis we can do and there are so many researchers can be performed on top of this data for sure therapy accelerometer data from the phone uh I think so yeah yeah definitely exactly exactly yeah sure any other questions object the question was okay the question was if we have some kind of like error estimation on top of it so it's really hard to have this in place for the items as you said for the items that was nutrition information are publicly available it's pretty straightforward but this is not necessarily the case and the nutrition facts for like popular brands may change over time so users have entered McDonald's Burger two years ago but it's different now from two years ago so that's really challenging but we try to use as much user signals as possible in order to evaluate so for instance if we keep presenting this food to the user and it's not selected a lot frequently we realize that this is not probably a good quality so we can probably take take this out from the verified foodies any other question all right thank you so much