Devreal

With Big Data Comes Big Responsibility

Event: Data by the Bay

data.bythebay.io: Nish Bhat, With Big Data Comes Big Responsibility

Recording: data.bythebay.io: Nish Bhat, With Big Data Comes Big Responsibility

I'm Nish and I'm the founding engineer of color genomics uh where we provide a physician order genetic test for uh the risk of getting various types of cancer and historically it's been very difficult to get this kind of testing done so we focus on increasing access in a responsible way and what I mean by that is working with regulation um and providing uh genetic counseling with every result that we return um uh and our test cost $250 um and prior to us getting a similar kind of test done could cost up to $4,000 um so that's me and what the company's about um and I'm kind of curious to uh get to know who I'm talking to a little bit um and uh maybe by a quick show of hands how many people in the audience I'm curious are you guys data scientist raise your hand if you're like a data scientist type Ro cool most of you that's cool um what about like would you classify how many of you would classify yourself as a bioinformatics person and there might be some overlap there cool a handful um data infrastructure Engineers or software Engineers or something few of those um anything else I missed just raise your hand if you're something else completely all right cool uh and all right so just to that helps me kind of level set get a sense of who I'm talking to um um uh for uh for most of you I I believe you all have the power to uh Advance the field of genomics and I'm hopefully going to convince you of that and show you uh show you how you can do that um but first uh let's uh talk about how things are done today uh so the field of genomics is quite nent out of Academia um uh the the way problems are solved in an academic setting where publishing kind of drives a lot of the incentives is very different from uh from industry and uh that the shift has been relatively slow to happen in the field of uh biology and in genomics but it is starting to happen and uh it's making things very interesting uh so I'm going to point out a few ways in which uh the genomics industry has room to grow and how you can help get us there and also if you'll humor me I um uh I believe bioinformatics has many parallels to the field of computer security so along the way I'm going to point out some what some of those similarities are um and uh hopefully uh Inspire and show you what's possible by drawing a parallel to uh how this has been done in a different field uh so let's talk about let's let's get to the first um first theme which is awareness and access so um awareness of genetic testing and access to it is small today but it's starting to grow um genetics is a powerful tool for assessing one's risk of getting a hereditary disease but but the standard of care is to actually look at the family history to uh see what one's risk is of getting a hereditary disease um now at best uh this is a if you just use the family history this is at best a proxy for the patient's actual risk which is uh which is largely influenced by their genetics um so the patient may have inherited a gene that drove uh their relatives cancer for example or they may not have uh conversely the relatives cancer might not even have been driven by a genetic predisposition um they may have even had a highrisk mutation but the effect of it may have been masked uh in the family history so a common example of this is a woman who is at risk of getting breast cancer because she carries a brca1 mutation which she inherited from her father and because her father was male the M the disease didn't show up in the family history um so for the for such people these people might might be at risk or they might not and they have no way of looking at that um without actually looking at the DNA um so today for many people um this is pretty much the standard of uh the standard of care in terms of assessing one's risk of getting disease um so clearly using uh DNA as the primary source of uh risk information is inevitable and things are starting to look promising you know perhaps this chart's pretty familiar to you um the cost of genetic sequencing has dropped over time at a very rapid clip and you notice that the graph is actually a semi log plot uh so Mo's law which is actually an exponential law is a straight line on this plot um and uh the cost of genetic sequencing is dropping actually much faster than that uh so um uh so this this shift coincides with the development of Next Generation sequencing techniques uh which I'll go into a bit more detail later on um additionally public awareness of genetic testing is increasing thanks in part to legislation um as well as public discourse uh more than any scientific advancement though uh it's been actually Angelina Jolie's decision to get tested for the brca 1 and 2 mutations that has driven awareness and demand for this type of genetic testing in a way that's uh to be more broadly accessible and since then a national and Global conversation has been kicked off that has significantly heightened the public profile of genetic testing as a viable option for people there is significant momentum for a sh in how genetic testing is done making now an ideal time to have a large impact on how it will be done in the future um so to draw a parallel to the computer security world uh let's go back to 2010 and uh so this is uh this is the facebook.com homepage and the thing to notice about this screenshot which uh this isn't my profile I took it from the Internet is uh uh the URL bar so back in uh back in 2010 everyone knew that SSL and https was is something you should be using if you're a website operator um but very few people did citing uh performance you know kind of vague performance um uh complaints as well as uh as well as complexity of implementation um so this is the way things were for a while um and uh then there was a browser extension called fire sheep that was announced and this is what it looked like basically you'd have a your Firefox window open and uh you'd be able to open up this uh this sidebar and basically it shows you social media profiles of people who uh are uh logged into these websites on the same wireless network as you um so you'd literally be able to see who which of your neighbors is logged into Facebook into Google whatever it is um and this was always possible it's just that this browser extension made it very easy to see how to to see uh to see what was going on um so all you'd have to do with this extension is just click on one of these and you'd be logged in as them um so obviously this was a little bit surprising to some people but nothing had changed in terms of the technology this was a pure awareness play um and it worked so very soon after this pretty simple hack was released um uh Twitter as well as Facebook made uh uh at least enabled uh https by default and now uh we've kind of come to accept it it's about 6 years later and uh I think most people would be pretty surprised if uh their bank or any major service they use on the internet that uh did not use htps um so I think uh so I think the awareness of uh of genomics is kind of following a similar uh parallel path and we're seeing that happen right now uh the next theme I wanted to go through is data sharing so it's currently the case that many companies treat human variation data as a trade secret uh there are even some public companies that use this to this data to partially justify their valuation um however uh as time goes on and everyone will get access to cheap genetic sequencing uh this proprietary information on variant frequencies will no longer be special and it's just a matter of time until that happens um by hold so by holding on to this data these organizations are uh delaying the progress of scientific advancement and the accessibility of action actionable clinical testing uh and this is kind of where the title of my talk comes from uh if a company's collecting large amounts of data they kind of have a they have a responsibility to use it in a way that's actually helpful for Humanity and not holding us back or being harmful um thankfully there's a huge movement across companies that are typically competitors pushing to collaborate and share this data um this slide is just an example of the types of data that I'm talking about variant frequencies as well as specific uh mutations um and uh many of these so many companies are pushing to share this data um and contribute to public databases such as clinvar uh of which uh here's a quick screenshot of that um and in the end uh the data should not be a competitive differentiator but it should rather be used to help uh help everyone move forward together um the real differentiator should be the the quality of the the test or the service that the company is providing um and the data itself as I said should serve to move Humanity's understanding forward and uh to draw the security comparison again um uh when these days when a security vulnerability is found that puts the Internet safety at risk um there's this organ there's a set of organizations that um make it so make it very easy to for people to collaborate across companies to make everyone safe um so and there you we've even seen marketing efforts around vulnerabilities uh that uh make sure this information gets uh gets to people very quickly um and uh to highlight two programs that are helping with this there's uh Facebook's thread exchange which is a program where companies can share information on malware and uh security threats with each other very seamlessly um as well as in February uh President Obama signed an executive order encouraging companies to share more data on security threats um genomics is seeing a very similar shift and I think the uh kind of the Precision medicine initiative ties into this very closely um and uh organizations are being pushed to share information uh more broadly so next I'm going to talk about software and data formats and how all of you can help um if we are uh if we are to generate useful data to push our understanding of Science and genomics forward we we're going to need performance and maintainable tools for uh analysis of data at Large Scale much of the software used for genomics and most bioinformatics software in general um has its roots in Academia so a typical development trajectory is that a lab or a grad student might build uh useful data analysis tool and even open- Source it but um and this tool will get dropped in as is in many uh genetic testing companies and uh but because these companies don't have a strong culture of engineering or open source uh these uh these tools won't be contributed back to um and uh this is it doesn't always happen this way but it it is a very common occurrence and uh as a result the project gets maintained either by that single grad student or by that lab or by nobody if the grad student graduates and there's nobody left to maintain this tool um so um so there have been exceptions but as I said but this model is still very common um and I believe the solution here is for software companies to or rather for genomics and uh medical companies to look less like this and a little bit more like these um what I mean by that is that they need to have strong engineering cultures that Foster development of high quality open source tools um and in the same way that company the companies like the ones listed here um spend resources building and maintaining open source tools that uh really the whole internet relies on um companies in uh in our space in the genetic space will also need to spend time giving back to these tools and improving it to make make them better for everyone um some companies in the space are starting to do this but we will have a ways to go until this practice is commonplace and uh I would like to say not all is bad and there have been a few really good examples of things being done right uh now let's dive into some of these success stories uh in bioinformatics and specifically let's focus on the problem of sequence alignment um and uh some of you may already be pretty familiar with this but some some of you may not be so um I'll go do a quick overview so um to cover really quickly how uh Next Generation DNA sequencing is done basically you have your DNA and it gets uh sheared into to many small pieces and these small pieces uh basically get sequenced by um by a sequencer it's basically a machine that you purchase um and the output of this um of the sequencer is are these basically these small reads of like 150 base pairs um so a very common uh yeah so the raw data that you get essentially looks like this um basically uh it's this kind of text format um and it includes the that thing in the blue is the actual sequence data and um uh the red is the rest of the read data um so to uh to talk about the alignment problem basically you're trying to figure out where in the genome these things go and uh we we can talk about the algorithm but we don't really have time for that so instead we'll talk about the format used to hold this data um and uh the first one is called the Sam format and as file says here uh it's an asky text format with very long lines uh which it is that's what it looks like um it's a text format it's not the friendliest format to use but the text is there if you need to debug something or figure out what's going on um so this is this was the first version and they also released a companion version called uh bam or which is a binary version of the text format and basically what this is is it's the same as that format I just showed you but they've taken blocks of it and they've uh gzipped the blocks so it's in this um bgz format or block gzip format um and uh the advantage of this it sounds like a pretty small change but the advantage is one it's a lot smaller obviously but two uh because it's um done in blocks you can actually index it and uh when you when you're able to index it you're able to have visualization tools like this that let you kind of jump to any region in the genome and see what the alignment looks like in O of one time instead of O of n time um and uh so uh the the more subtle thing that this allows for is for doing this kind of operation over the Internet so you could imagine a client in a server where the server is uh serving both the alignment as well as the index and all you need to do is fetch the index and you can do an HTTP range request for just the data that you need so you're not downloading the whole maybe 10 gigabytes of data you're only downloading 16 kilobytes and there's a tool called IG VJs and here's a quick example of uh that live being used and the main thing to there's a lot going on here but the main thing to look at is uh those 2006 requests at the bottom uh and the 64 kilobyte response sizes that show that we're only fetching the data that we need um so uh uh so uh adom is another uh is another tool that's used that's kind of going A Step Beyond the uh bam format of sequence alignment um we've uh heard it come up in a few talks at this conference which is great um and uh B it basically allows you to query query your alignments in a way that uh the BAM format just isn't very amendable to so uh it stores data in a columnar format so you can make queries across all across all of the reads in a in a position whereas with a bam format that might be prohibitively expensive um and another another step forward is Google genomics which has been kind of putting offering a nice rest interface to this kind of data as well um and all we did here I should note is compress and index the original format and then move it behind an API so these this is all the kind of stuff that you all are familiar with doing and um there's still a lot of lwh hanging fruit in terms of uh in terms of further progress to be made um and uh to draw the security analogy really quickly I was going to compare it to to uh pgp which uh you know this is it's this tool that kind of came out of an academic setting but it's still used um but because so many uh smart software Engineers have been thrown at the problem of uh making encryption you usable um it has become more usable both for uh this an example of how it's not usable but both for for developers as well as uh for end users um endtoend encryption has become much simpler and uh this is a quick example of um of one place in which we have yet to improve which is the variant call format so um we're still storing our genetic variance in a format that's text based so there's definitely a lot of uh lot of room to be improved here as well so uh to conclude it's an exciting time for genomics the technology and our ability to read DNA is getting better by the day but analyzing the data remains a strong challenge um and uh uh data initiatives to get people to tested as well as to share data are lending lots of momentum uh to this to this field which is making it an exciting time to enter now um I'll close with this quote the best way to complain is to create um and wanted to say thanks to uh to data by the bay but also to all of you guys for giving me your attention and time um and if any of what I've talked about sounds interesting my company color genomics is hiring um or just feel free to get in touch with me and talk about this stuff all right thank you very [Applause] much I uh I think we've got a h head for the next talk at the moment but hopefully Nish will be around for sure yeah happy to answer questions in person fantastic oh cool okay great all right all right well I'll be around afterwards too great thank you so much