Devreal

A Scalable GA4GH Server Implementation

Event: Data by the Bay

data.bythebay.io: Andy Petrella, A Scalable GA4GH Server Implementation

Recording: data.bythebay.io: Andy Petrella, A Scalable GA4GH Server Implementation

yeah thank you I didn't say anything yet so just little bit before um okay thanks everyone to be here uh it's really a pleasure to come at De b um to talk about what we do uh in agel science and specifically dedicated to to gen mix Information Systems right uh so right now I have two talks in a row uh the first one will be about interruptibility of of systems in uh in genomic specifically we'll talk about something called Global analytics for genomics and health I will give a few words about it um but before that so DEA FAS is actually yeah it's a Belgian company so far we are cting right now now the entity in the US and we expand our Market over here and what we do is that we enable data driven businesses by allowing data sentes to be more you know productive thanks to agility and the toolkit that we are creating is is is allowing them and is on right hand side you have a shortcuts but the main topic that we are doing is data science on data science um so I'm in Andy Petrella I used to be the gorilla um I have mathematical background I did a lot just spatial data analysis in my in my uh professional uh life let's say um I created my own disability Computing 15 years ago to do this kind of things um I created a spy notebook uh who knows Park great who knows a spine notebook okay cool I mean yeah um so yeah I'm the trainer for Scala spark I gave a training with oilly on Pipeline a pipeline a few months ago and my buddy zav to with whom I created the data FAS is um has kind of same background in physics B informatics he created also his H Computing while ago he's a pearl guy initially genomic is about pear actually so um and yeah we both did a lot of marchal learning so okay the lineup is g4h G G4 GH scalability a bit of infr architecture what is a catalog and the extensions that we might foresee so what is J J for GH right so genomics for aliens for genomics and health okay so the thing is so uh it's a kind of framework if you like it's like w3c but for genomics and health right so they are there to to create standards on on the data and and and so on but not only genomics but also clinical data why the high behind that is that we want to uh make advances faster in genomics and health right in order to do that we need to be a able to share information in in a efficient manner and in a clean manner as well I don't mean that the data is clean but the the way you will share information is at least clean how we do that okay so we Define inter reparable services and types um I know I don't know if you guys are are really familiar with geospatial but these kind of things exist for 30 years in in geospatial right so now it's time to have it as well in in in bioinformatics and gemics and health so who is G for G it's 300 plus institutions mixing research Healthcare um also very big Enterprises like Google Amazon are there right in order to uh move forward a little bit faster with their uh footprint okay so there are a few working groups out there um clinical which is more dedicated to the data itself so what kind of data we can share the data group is about how we can represent the data and share it Regulatory and ethics they are there to say okay what we can share and what we cannot and in which ways if necessary security is the the protocols mainly on protocols if we can share this information what are the constraint and what we have to respect in order to be able to share it so this is not this is not I mean they're not joking right so they're taking it seriously uh why it would be that important to have something like that I mean the data available right now is huge so we are now at more than 100K jobs so that was last year last year we target 1 million and very soon it's going to be 10 million and of course we target the 10 billion we cannot really go beyond that because we're not that much uh but I mean we this is something that will still growing and I'm still talking about the genomes I mean the human genomes there right so the many genomes that you might also uh take like environment bio and whatever right uh but the idea is that what if we could carry this million of inform of genomes which are pretty huge right um what could we discover based on this you know huge information that we have in vable right now um and the idea is that not only the genomes are important as I said but also we want to uh you know mix it and and Mudge it with many other data set like imaginary pet scans and thing like that um also being able to include real time like monitoring okay so how we can correlate both things and and of course one important thing right now with smart cities and so on you know environmental data can be something very interesting to mix up you know how to see if there is a collision for instance between uh genos mutation and environment um uh changes for instance who knows right um so yeah so in order to be able to do that then we need to be able to carate data and to make sense out out of it right so we need to be able to quer it at least so not only one data set but all of them at once for instance in order you know to be able to to query so we need to be able to know which query want to issue that we need to be able to discover it right and this is a very difference between discovering and quering right when you query you know what you want to get back more or less when you discover you you shoot in the in the in the shadow right so you need data to be distributed and we need of course to have systems to be scalable wa yeah okay so um what is about disb systems so you know all spark right so we don't want to copy the data we want to work on partitions uh and we want to follow the protocols but of course we want to do that without any complexity or additional complexity who works with HPC okay so you know how great is is how easy it is you know to manage to schedule and so on so it's a very um hard to work with this environment now spark um put some some sense a layer of um of I mean facility or it's more easy now at least to it's easier now to to work with these kind of systems right so I'm going to give a short information about Adam who knows Adam here okay a few of you guys so Adam is a um layer for genomic information right so Adam is uh is part of the Big Data genomics project in amplab and it's composed of many things this uh this uh uh big data genomics that you have ago for the data sets avocado for the VAR color Adam for the core API and the andio so how to be able to carry the data and how to store it right and bdg formats is one of the most important one those guys are the the types the schema right dedicated to genomics and the thing is that so we need schemas right you know we need schemas because who for who knows VCF jcf bam and Sam and so on those those formats are not suitable for for disability Computing why because there are big data sets it's like a big image right so you cannot really easily um compute things on a big image because it require a lot of shuffling you need to share a lot information in order to be able to do that and we need to be able to rethink the way we are storing the data or producing the data right so that it can be stored in a efficient manner in a distributed uh distributed system like hdfs or S3 or whatever right and to do that so there are many maners to do that but they took in Adam the AO um protocol let's say right so in the in a you can Define uh schemas so you define the overall information uh that you want to share and and behind that you don't Define the the storage but it can be distributed but the thing is that what is really cool is that not only you can Define the schema for the for the information in and out but also you can Define schemas for the methods so who work with soap right so arrow is like soap is a beef up soap if you like right so soap without XML okay so you are able you target a system you say okay what kind of operations can you do okay I want this operation what types do you take what type what types will you uh send uh to me okay so this is essentially what it can do and it's part of the protocol as well it's not schemas that are outside like in soap so it's part of the protocol right so as I Said So Adam defin strict schema and um you can store the way you like it behind that uh and the idea is to store it in a uh partitioned way so um so you can provide you know smart interfaces and so on so forth but the interesting thing is that in Big Data genomics in amplab uh which is that been R named now rise or something I can remember the exact name anyway cool um so the thing is that they are not on top of that creating a few you know martial learning algorithms or you know reg how to say that real welln uh analyzis like MCT uh some realization as well they do very good at that variant calling so it's necessary in order to exate information over there the novo assembly and so on so so they are working now not only in the so they they are on one level H right so they Define everything uh below and then now they are able to do very awesome things with that so in the next talk I will give a little bit more information about that so I would have liked to swap them but yeah it's hard at the last minute right so scabber server now so what about you know we describe Adam and now we get back to gf4 GH so they are there to share information efficiently to provide the right methods and so on so how could we do that since we have this amount of information available and want to share it efficiently um it's better now to maybe have a server which is schema oriented what why because then you can ask um your your laboratory and you and you ask another laboratory what kind of data it can have and how they can C it right and it doesn't have to to uh share how to say they don't have to take a text file back and then parse it in in its own way right you just take the binary format and the the protocol handle it right and he has regular types in his hands so there is no uh how to say um communication involved in here and certainly not human um okay so it defines formats uh for the data but also for queries so this is something looking like Adam but in a more interoperable way okay um so they have a multitude uh of services one for the reads which is the the small information that we can have in the genomics in the genome so when you read information the bio informations uh the variation uh the variant sorry um and also it's not there the genotypes but anyway and beon is something more for metadata right so I'm proposing just in small architecture uh something that we have implemented in DEA FAS uh it's still open source right now uh but it's mainly you know able to to implement this protocols in a very scalable way thanks to spark and Adam right because those guys you know j4 JH they have implemented this this these services but using python you know local scripts and so on so it doesn't scale really well and now what what we propose is that using spark instead of Adam for instance in order to do that because then not only you will be able to store the data efficiently or to compute it efficiently uh but also you will be able to use the new algorithm available in MLB or h2o or whatever the whatever uh libraries that are coming right now in the uh area of data science let's say so you you can have this kind of microservices orientation so maybe one service is targeting only one method or are you Gathering the method together in a more semantic fashion but the idea is that you can have multiple microservices serving one purpose and then you can expose there your your a RPC which is like soap RPC right so just a port which is open where you can Target um you can issue a query like what can kind of operations do you have and then the pro protocol goes down um but the thing is that you can also expose it in a rest API if you like it right then you enter a uncertain area I would say uh but you can add autentication UI and so on on top of that but the main thing is that on the back hand you can have a you can use the data source API from spark to do that we can use beam or whatever you like in order to access the underlying system that holds data right and if you can if you don't want to move the data from the old storage you know the silo storage where they have terabytes of data and and and you think it can be efficient then you can have a protocol from the fpt layer that goes through the data layer and and somehow you can cash it for instance in ton for a little while right so you have a small process that will hle the data and push it into a cache like I I I still think about taken like a cash I kind of like this uh to to think the the this way actually and yeah and if you want to have a persistent cach then you can use the underlying pers system capabilities of ton to store it in C whatever so but I mean this is successful um uh implementation so far uh we'd like to also push it further um but yeah this is essentially what we propose so our implementation is very scalable so we tried with four noes 40 noes to do some analysis the thing that we'll talk just right after um but the thing is that the processing is mainly based on spark and Adam but also we are enabling extensions or custom queries now okay so okay so the standards are okay so everybody can use them and it it satisfies like 85 person of the of the use cases but I have also internal models so I have internal processes that I would like to expose and in order to do that so you might want to add an extra micros service that will do that so you expose your service and every schema and then all Laboratories can hit the server not the operation not the types not the outputs and there is no communication involved right so you can expose your own stuff custom CES to anybody in no time so you don't have to share anything um uh and then you can plug in data sources I has shown so right right now we uh we have implemented what was very important for us which is variant information we didn't go to the read uh yet but we have implemented like half of the standard in you know in in a few weeks at that time um however so now we have Services everywhere right um but it reminds me the time where internet was you know ftps and you had to know exactly the address of what you wanted to consume right so it's pretty much like that now you have Laboratories you have institution everywhere in the world that means that how the heck could you know where the information is the where the data is that want to consume for your own purpose right so now you need to be able to discover not only the data but also the the services so you need to be able to catalog this information right data and services um so the the idea is how do you know where to query um which custom queries can be available and what I which ones are more important for me than the others and then as I said so we need a search engine for this kind of services like Google was you know or tamahu and and other services um alav Vista as well for for the olders um so that's why we we have this genomics catalog and the idea is to extend gf4 G with extra capabilities uh that allows to support metadata queries that means that a schema is describing service capabilities but also it's also describing data set that means that you should be able to Target to to issue query to a catalog saying what kind of services right now where is deployed about variance for this population in this range okay so you should be able to do that and I'm not targeting each institution and and once one at a time but Target this catalog and they say okay do Services there there and there contains the right information apperly so you can then issue the right query in the in the right manner right and they may have different version of the of the of the services that you don't care because the protocol is handled handling that right so the functioning is in there as well um so the order in order to do that you need to have harvesting operations that means that you go from one service and then Services can reference each other and then you crawling the web of services this way so it is pretty much like that it's done in just spatial for instance you know so the data is always connected so there is a way to do to do crawling in there um and finally yeah so Define a few queries that allow to search efficiently the data and the services into into into this catalog this has to be also uh very um you know standardized so the approach to do that is that you know in order to be able to do this custom queries what we do in deas is I said so we want to to enable productivity so instead of having the the data scientist doing his work and then converting his work into a project and a service and AO and so and so on we do it for him that means that from The Notebook we compute the job we create the AO schema for the types for the Cassandra layer uh table or whatever hdfs information they have stored and also we are building this knowledge base in order to be able to query information and service internally so it looks like that actually I will go through because I just have 20 30 seconds so we have the notebook and then we click on deploy so you define a few things like I would like to name this like that what is the the the version then you deploys the project uh this this is deploying a project a job in spark like mesas uh Kronos for instance this is to define the service then you end up with two Services because they were a cmra table so uh you will be able to expose a model but also the the the the information and then we have integrated this search panel right now that where you want you can search what kind of data you looking like and then it returns jobs services and so on but now the idea of data science on data science is that you don't need to use this thing anymore because while you are typing your work in the data book in know notebook then we will be able to discover what you want to use and maybe proposing maybe this data set or maybe this service or whatever so this is reversing the way you will um use catalogs and and and data L for instance when you click on a service then it loads the schema and then you know what you can do so yeah so I'm I'm done actually so this is uh a article did you want to to look where sharing is important we need you of course to do that and to to go beyond what we have now and this is pretty much it so it's a genomic on steroids right thank you qualification people looking for this is no this this is me 10 years ago actually thank [Applause] you if there is any question I'm open to yeah uh one thought U if I look at your catalog yeah on as well I think prob under uh deliver what maybe potential you can you have because if I have a I want to do I probably don't want to bother even go to look for the service I would rather just say do the job for me it's a best way I can do which job which kind of job then if I have General related things I want to do and with what you have there it's best for me not to worry about what service I should look for set exactly yeah yeah this is the ID so we can send the service but then I mean this is your task to decide if this service is very accurate for you or no and when you consume the service you might want to know okay is is the data accurate enough for me to consume for my own project so what you can do then is that because we are doing that in our product is that you want to maybe go one step deeper and see what were the sources that were consume in order to produce this data set and then this concept on line AG important though but that's server is say this is good source it for me and make yeah sure sure you can do it but then it's your decision to use it or no now another I have been feeling [Music] something yes will help for me more like they have a very good a good protocol which one yeah yeah yeah and this really addressing this not actually yeah for for in the protocol yes exactly yeah and so from try to address that's actually no it is really important because just stay there and watch the next talk then Okay because I'm going going to talk about it right now okay about the schema importance and yeah but it's a good it's a good um uh question so I'm going to just talk about it right now