Devreal

Interactive Machine Learning on Genomics...

Event: Data by the Bay

data.bythebay.io: Andy Petrella, Interactive Machine Learning on Genomics Data with Spark Notebook

Recording: data.bythebay.io: Andy Petrella, Interactive Machine Learning on Genomics Data with Spark Notebook

welcome I come to my son go talk for who we're mad enough to stay and to the newcomers you would understand so yeah so let's it's a twenty ten minutes talk so I would try to be as quick as possible regarding what I want to say the main thing is that i want to evangelize you know the way we can work with genomics today in an interactive manner using something like expanded book right so in a few words about data phyllis diller fellows is a Belgian company so far but we're creating us so the incorporation right now in the US we have been accepted into the alchemist accelerator program which is the kind of Y Combinator for Enterprise editions and what we do essentially is that we we aren't visualizing the way to do della science and prediction that means that we are you know adding the agile methodology in the system but at the tooling level not at the fizzle fizzle outside of this difficult layer and based on this toolkit too so we are really interesting what we do in on data science on data science in order to be able to go from the assistance to the data scientist instead of the descent is going to the data leg for instance so I am the gorilla my confounder is this dilemma because Lamar cool right they are definitely cool right and so we both have a scientific background in math from my side physics for folks I VA we we've done a lot of things we created our own disability computing in just patient and genomics 15 or maybe 20 years ago for him he's older than me so i created this by notebook who knows park here okay who knows the spine notebook okay good good learner yeah and we are training this guy is a pearl programmer and and we've booked on a machine learning of course so the Lana for today very quickly so I'm going to talk about genomics data what it is why we need to speed data sets also what kind of data represent a storage we might need for this information we all take a look at a dynamic use case pretty common let's say what we can do with this interactive computation to target this side into to solve this use case and finally what kind of machine learning libraries you may use with genomics today but but that you might not have envisioned so far so dynamics that are who is familiar with genomics information here okay so i can i can give a few few information about it really quickly so first thing is that so it's very fast in production right now so we are very ramping up and the number of genomes that we are producing every day right and i'm not only talking about you know human genomes but everything flowers I mean plans or I mean inside so whatever so we are sequencing a lot of things and this is really ramping up Francis for if you think about a human genomes is like three billion bases per sample right and for each sample we are at least doing 30 to 60 take thanks okay and that means that we have a considerable amount of data over here right for one single people web salesperson and the idea is that this information is produced not in a very linear fashion so we don't have the food sequin at once we have this information split into reads which are more or less 150 base / / rates okay so this is already distributed right so in the biology in the chemical layer of this process they are already distributing the computation let say or at least the the measure and the thing is that only producing the data say is that as it is by itself very important very I mean resource consuming right as I said so we do 30 to 60 we hope take on the text on the on the sequence that means that we want to be able to have enough say overlap between the reeds in order to be able to you know reckon strike one sequence with a with a higher accuracy so that means that from this you know this mass of ivory is there we reorganize them into a linear fashion and then we compute you know statistics in order to be able to say okay for dispositions probably this one up maybe this one with this probability then we have what we call very variable a variant sorry and this is the variable position in a population let's say right this is the information that we generally use because I mean this is where the variation is right so we can strip down the information to this variance only and there are 1000 of them right so from the trim le billion you have only 23 million for instance of 300 million sorry variance so it's specifically it's a genotype which is composed of a very on which for the post for the chromosome at which position this position we have the alleys let's say a or C for both sides and then the genotypes dereferencing one sample and then we have information for this variant which will be either I see depending which side you are places but the thing is now this variance we have tens of million in general and and also for the samples there we are reaching very very high number so maybe 10 years ago we had like 1,000 samples available right and now we are reaching the million so this is where we are going right if you think about what's going on us there is is 1 million genomes project has been started by by the President Obama so imagine 1 million of this disarmament information this is huge right and youthful not used for nothing right this is there is a high opportunity there so that means that we have a lot of stricter data so I want to invest in the fact that we have structured data so we know exactly what we have okay yeah I mean exactly the thing is even though we know what we have the description of what we have is done in PDFs because information is stored as like RFC's right so we have this formation is described like that so that means that if you want to know what's going on with the data so then you have to read this paper interpret it and then implement it right with a lot of interpretations as i said already so that means that these things is pretty much okay for human right for humans but I mean when you want to go beyond that when you want to have interaction between machines then you're not in the right area right so in order to go from one side to the other are going to do the notion of schema so we extract the exact information from this de spdif right and we create a exact schema that we can reuse share and and if make evolving in a in a let's say [Music] sustainable manner right with different versions and so on right so and schemas are very good for machines of course because this is a protocol it can be and can be respected by constraints and rules and a schema looks like that so it's an average schema in this case so we have a record and record is composed of several fields that might be older records like variants or we can have also an array of alleles others there and and so on so we can have nurse as well which is bad but necessary but anyway so we have no types and this type comes from Adam and Adam is is I think I have a few slide about it afterwards but Adam is is a project from the Berkeley and blab and they have defined actually these types for everything related to genomics so far right reads variant whatever but also the very important / tons of average that they are not defining yet a storage so you are able to d2 to plug in your storage the way you like it right but a good way to do it of course is to use something like parque y because parka is distributed right it is schema based as well so it has there is a strong correlation between the ski mine avro in parque and also this is the most important one for me you can read acquire it efficiently this is one thing which is very important that you can read a query efficiently you know we using types and also is pretty much compact because it's partition and the partition are compressed right and there isn't a compression algorithm which is very efficient because actually Parker is culinary anted so they are able to compress the data like crazy i wade through it's very efficient why you want to read the data is because General William on you when you do genomics analysis you will consume ranges of positions into your into your data set that means that you will consume a full data set but just only one information into the data set that means that you since you have columns next to the other and information for one column next to the other on the distance you would you just scan read the disk and you have the information right away so it's very very efficient so you can load multiple items and and then thanks to that you can even load them multiple times which is not really easy when you have a big charge ezip stored because you don't want to lose them any too much storage right so you cannot really easily recreate and to read it multiple times so we worked a lot with the 11 genomes project why because the data is open open source the data is put on s3 as big busy to file right so they have this project is that run from UK and they have sample many people from different area in the world and they condense everything into Cambridge somewhere summer they did it in in Cambridge and the end but the information is very available open source so you can consume the way you like it it's on s3 as I said as busy to file the thing is that since you have busy to file right you have to download it not first you have to find it because to be honest the folder is a mess so of course you have the for his data but then they made some transformation and then some analysis on that then you have a mess of folders and then you have to find your own information and more or less moreover there is also some Delta increments so they are adding new data sets in different folders so it's very hard to find but anyway when you have it you download it the future bites too I think then you can zip it or you wanna be zip it so it takes quite a long long time hours actually when you have it on come uncompressed you you can load it then you can convert it in Adam for instance you save it and then you can load it in one minute okay so from ours using the right tools right you can load it in one minute the same data set okay and Reagan the size we have a raw file for part of information which is 152 gigs right and it's compressed with the same information but clear right we have several gigs we have reduced by to the the size but raw data is not readable the atom data is readable like they're like this I mean you don't have to do anything also we had 33 three parts for the raw data of course one per chromosome and for Adam we for our tests we recreated like ten thousands partitions that mean if you want to use that in this way the matter you can reach a factor of prioritization over 9006 you like so this is prezi amazing win and and also the raw data per say when you haven't compressed it is not even easily distributable because the schema is on the schema de the headers is of course in the header so you cannot share information efficiently and and the PDF the lights on before are not very strict so they are open to some modification so it's very very painful but to load it in one minute so what you can do is that you can just use park right it's parquet right it is very good at that so you you take your spec contacts you inject the atom libraries and then you say Adam load you pass the file to that which is on the street and you have it then you return what a GD of genotypes it's typed is not like strings so if you had read them or some files you would had byte byte strings or byte arrays or you would have just trained right right now you have a genotype so you can call functions on at least fields so this is the good thing so now from from the notebook like there you can start consuming the data and starting your expiratory faces the problem is that actually we have quite a number of samples right so the data is starting to get pretty huge and if you want to know is that if the data is it erogenous so if you want to see if the if the the samples that you have are part of the same let's say population right how to know that right in order to know that so you will start to do some clustering right so without the first the first principle that means that you want to you don't want to model the way structures are created into this population you want to just learn from the information available right now so you will do in supervised learning so most of you know spark so the interest of spark here is that you can scale you have a few layers on top of spark which are pretty interesting from manipulating Dale I like graph processing this is my my preferred to be honest much learning a little bit as well you have not only Emily but you have older ones and you can also optimize the way the memory is used and so on so forth right spark spy notebook as I said so we want to be able to do interactive manipulation of our genomes why it's easy actually because you don't know what you will do with the data you start looking at it right so you want to be able to do manipulation of the dayla many many times before you just found what you want to do and again right so you have to be interactive and not launch jobs and waiting for it to come back with information like you know you can really do unit testing when you do data science right it's really hard because you don't know what you want to do to go you can do it afterwards but not before so this my notebook is interesting for that because its interactive of course it's cannabis so you will be able to use the power of the types it has a reactive and playable charts API so you can have your model train with the fancy chart showing how oh good they are you can only stop them it's easy to install because it's only a HR gzip with no dependencies on the system and the thing is that you can have multiple people using it with on conflict because it has multiple cysts park contacts so each time you launch a notebook you have a new spark context that mean one guy can work with the atom of 16 and the other guy can work with the old the atom or 19 and there is no conflict because they are not sharing the same class path in the same jvm actually so they are new JVM spawned each time it's interesting because then you can have this huge function right there in a Cell right and it's doing very fancy things right it's pretty unreadable but the thing is that you can split it right you can split it in such a way that you can check with the types what's going on well or wrong write it smelling manipulating manipulating peppers they're right and you can very easily doing errors but thanks to the type you can split it and one time I see okay I control enter and I see what time i had type type i have and then i can be hired by the compiler to be able to move forward and this is the importance if we get back to a i mean a dynamic language and we had met this efforts to have the types for genomic information that we will be losing quite a lot of advantages of course of the of the schemas on the data so yeah this how you would compute frequencies for the Allens right so then you can compute it automatically the notebook detect that it's a list of double and int and then it prints a bunch information the other plots you can do rather you can do parallel coordinates if you like whatever but i mean this is the automatic ones and you see the distribution that means that most of the variants are shared and then are pretty old sorry and then the variants are is this good for instability is decreasing this is machine learning what is important to be distributed when when when it sees is as fast as possible as quick as possible as soon as possible as well so why you want to have this machine learning is because our data sets are getting huge mirchi terabytes we are reaching the petabytes in this area we have huge models right so you have a teeny cpu model in one machine and we want to have distributed algorithms on that and also we want to do very large assemble so we want to run many many models and assemble and one time that means you can use the distribution environment to do that you want interactivity on top of that otherwise you might take a lot of time before being able to do something great and at the end you want to have your results to be services service-oriented right so you don't expect a geneticist to to look at charts at the end so you want to see generally the raw data in order to be able to hike to on the system let's say to give the right medicine so it just is very happy with one number saying the probability and then of course with the stupid environment you can use the you know d you can optimize the way you will launch up job thanks to constraints on your face texture some spark and my lab is one good option you have these things I want going details to that but this is one example to do in supervised learning this is the classical Kimmy K means but we do it on genomics data right it's two lines you load the data you prepare it a little bit in order to have right matrix there be used I think we use their the Manhattan distance between positions so you have for for a very small space then you can compute Manhattan distance between points because you can call it ABC ctg a as four points in the space anyway but then you can very easily compute distance and then compute the k-means and at the end you can separate population pretty pretty I mean it does a pretty good job without any optimization here you have three populations and the clusters and you know the the matching are pretty good but the other one that you can do that so is h2o who knows that she were here right h2 I say is a very okay so I I think okay so i just have to slide i want to show h2o so is there something that you want to look at yeah because it hey I've implemented the whole memory MapReduce but it's very highly optimized for circus and Jamie M it looks like that then you can do it this way so from spire you can create your edge to a context and from there you can start using h2o I recommend you to watch this video is my buddy zavia that that does it that presented something it h2o word recently and mainly what you could be able to do is run at show context as there are frame on the data and then you were able to run some deep learning algorithms in two or three lines or a few more lines and it did a pretty good jobs as well right some takeaways so right data for the with a schema if possible work interactively prepare a big infrastructure for that good margin learnings libraries that you want to use right now and being a very close to production is very important so that you can deploy your work as a service efficiently that's it sorry for burning the time you