Devreal

A High Level Overview of Genomics in Per...

Event: Text by the Bay

sfspark.org: John St. John, A High Level Overview of Genomics in Personalized Medicine

Recording: sfspark.org: John St. John, A High Level Overview of Genomics in Personalized Medicine

right hi I'm John st. John I'm one of the cofounders of driver group we're a company that basically takes takes the tumor and the normal from late stage lung cancer patients and from that data we give patients a clinical report that tells them exactly what kind of drugs they would most benefit from so today I'm going to talk about what that data is how we collect it and what it means so in this talk today i hope to you know teach you all about that and also at the end i'll be showing a little promo slide we are hiring so if you are interested in this kind of stuff please talk to me so I'll start by talking about what cancer is at the molecular level I'll then talk about how we turn tumor just tumors raw tumor samples into text then we'll go into what the data means how it's actually help patients and then I'll go into some like next-generation analyses like actually modeling what's happening inside of a cell and what that looks like followed by opportunities for learning and getting into the field all right so cancer is a disease of normal cells in your body that just start going bad okay so normal cells in your body know how to act and you know they know how to behave they they don't try to grow uncontrollably they don't try to out-compete their neighbors their you know well behaved cells in tumors that's not the case they forget how to behave nicely and they start dividing like crazy and they start forming giant masses and spreading all over your body so basically it's like a little piece of you that is now trying to take over luckily for you there are a lot of things that have to go wrong before you get a tumor okay so these are the different the different things the different kinds of biological processes that actually need to change before you before you develop cancer so just as an example of focus on one of them which is avoiding immune destruction so the immune system works by you've got these cells called dendritic cells they just float around inside your body and wait for little bits of protein to float around and if it finds one of these little bits of protein and it is not a bit of protein that it recognizes remember tumors are filled with mutations that causes different kinds of proteins if it's not a kind of a protein that it recognizes it'll let a t-cell know that hey I found this thing floating around this is bad you should go look for it so now a t-cell if it's behaving the right way will go and it'll actually hook up with the tumor and it will destroy the tumor cell so your immune system is actually capable of destroying cancer if its operating functionally around the tumor so tumors have to avoid that again that's one example of quite a few different things that all need to go wrong before you actually develop a tumor that's going to kill you so Frank mentioned DNA it's huge something really crazy to think about cells are so small you can't even see them right they're like tiny tiny tiny things if you were to take all the DNA in your cell in just one cell and stretch it out it would be one and a half meters long and that's like the physical length of the thing DNA is composed of a's sees GS and T's represent little chemicals it's only ten atoms wide so it's a very very small thing it's kind of remarkable to think that we can even sequence this stuff and even see it this is pretty new technology it's you know maybe 10-15 years old that we've been able to do this really fast so it's an exciting time to be involved in biology right now so what goes on in a tumor I told you a lot of things need to change a lot of things need to go wrong before you can get a tumor well you typically accumulate mutations so you know it's you Julie not one mutation it's usually a large number of mutations what is a mutation it's typically a change in your DNA so this is representing a person's normal cell and a person's tumor cell and you can see right here there's a little chunk of DNA that has been inserted it's been added to the person's tumor that could potentially cause the tumor to start misbehaving in one of the ways that leads to cancer so I don't know how many of you have heard about the central dogma of biology it's this whole idea that you've got DNA your genome with in your genome little pieces of your genome in code RNA right so this is zooming in on a gene this is zooming in on a piece of DNA which encodes RNA which encodes a protein a protein is a little glob that floats around and actually does stuff inside of your body okay so really it's the proteins you know that you have to be interested in but as a limitation of current technology we can't look at proteins directly very easily in large numbers so we typically focus on DNA and RNA which are easy to read so what can go wrong what does a mutation do well as Frank mentioned the typical things that can happen you can have either an under-representation of a gene so not enough of a certain protein could be present and that could be bad in a cancer cell you could have too much of a gene so for example say a gene that's involved in you know developing embryo right so you could imagine that that would be bad if all the sudden you're expressing developing embryo genes as an adult right so that can happen in tumors you can also just have qualitatively different or just different proteins different shapes of proteins that now behave differently and maybe they don't interact with other proteins properly so that's what's represented down there so how do we turn this into text how do we turn this into something that we can actually analyze so this is DNA sequencing typically it will do and this is starting to change people are starting to look at the whole gene everything involving jeans and also everything between the jeans but typically what the medical field does today including our company is we just extract the parts of the genome that turn into proteins at some point so this is what we call the exone so instead of 3.2 billion characters we're down to 55 million that we have to look at so still quite a bit of data and then in addition to looking at the at the DNA from the from the exome we also look at what it turns into the RNA as a proxy for the protein so it gives us some idea of how much protein is being expressed in a Cell how much protein is present so these are sort of the two different kinds of data that we typically look at both of these kinds of data we can biological data we can feed them into sequencers and read them the same way at the end they create the same kinds of files so how much data is this you know why do we need why do we need spark right so this is one piece it's a really huge book I haven't read it I just know it's huge if you were to replace every single letter in this book with an AC g or t 14 hundred pages long three million letters you have to sequence a lot of these books okay so we're looking at about six thousand copies of war and peace because of all of the redundancy that we need in this data and unfortunately one limitation of modern technology is that you can't read the DNA end-to-end right you can't just read it you know position one to position a thousand first you throw it in the paper shredder so imagine taking like several bookshelves filled with books throw it in a paper shredder and then what does this represent well now you've got to find some mutations right what are mutations their little typos you have to find tiny typos maybe a hundred of them in this huge pile of data right you have to find slight ships and DNA abundance I'll talk about that in a minute and what that means slight differences in RNA abundance I'll talk about that as well but really it's like there's so much noise here and there's so much data here you have to have some pretty powerful methods and you know a bit of knowledge about how to really process this the right way so now let me introduce you to the data format that you start with it's called fast q there's absolutely nothing attached to it so none of the nice stuff in Frank's schema you don't know where it belongs you don't know where it comes from all you know is a unique identifier that says that this sequence has some kind of name and it's different from other sequences the DNA sequence itself typically about a hundred nucleotides long with modern technology and then confidence scores for each letter that's typically what you get out of out of modern modern sequencers so now I mentioned the first thing that we're looking for on the order of about a hundred maybe to a thousand mutations in a typical tumor I didn't mention before but we're actually looking specifically at lung cancer someone asked about how much you know aren't tumors really messed up well it varies quite a bit all right so lung cancer is actually one of the more messed up kinds of tumors lung tumors you typically have mutations all over the place it's one of the most mutated tumors so I guess we're lucky in a sense that we're analyzing lung cancer because we see a lot of these so we have a lot to work with so what we typically do is a comparison between a patient's normal and their tumor so that gives us this diff that Frank mentioned a few times right so what I'm showing you here is an actual mutation in a gene egfr thats related to related to lung cancer and if you have this mutation that means that you're eligible for a certain lung cancer therapy and it will extend those patients lives about an extra six months over standard chemo and radiation so it's a big deal if you can find this but you're looking through a lot of data to get there you have to have pretty sensitive methods that are also very specific you know false positives are huge when you're dealing with like 50 gigabytes of text data filled with errors so I also mentioned copy number changes chunks of your chromosome can get duplicated so remember earlier we were talking about genes and expression levels of genes well one way to get a higher expression level of a gene it's just put it there twice ok so this happens all the time and tumors literally a chunk of your chromosome which is you know part of your DNA will just copy itself and it'll just be there twice so how do you see that right I mean it's not actually like changing any letters anywhere right but what you can do is you can just count how much data you have in a given window in your tumor sample and in your normal sample and you can do a ratio and you can slide that window along the genome and you can look for significant shifts in the ratio that are consistent right so you can see there's tons of noise everywhere but you do see these very large very significant shifts so that's what copy number change looks like and then I mentioned RNA this one's tricky this is maybe one area that is most clinical labs don't even want to touch it it's really tough to deal with it's very noisy data in the expression level it's hard to interpret we are attempting to go there right now our lab is so there's a lot of work to be done but I'm hopeful there are certain kinds of things that we can do right now with it for example so we can classify what kind of tumor you have just based on the RNA data which is pretty neat just to see that kind of a biological result just fall out of this data but you know what is this data right so just like in copy number change what you can do is you can look at the RNA and you can count how many reads in the RNA match up with the tumor versus the normal so it's another ratio kind of problem right and if you plot it out every single gene and what it looked like and just highlighted the outliers you could look for certain genes where there's way more of it in the tumor than the normal okay if it was exactly identical it would be on this line right here but what we see instead or you know some genes that are high magnitude outliers that are really shifted where you've got a way higher abundance in the in the cancer sample than you do in the normal sample so those kinds of things are interesting as well as these you know really low abundance abundance rnase so how's this actually helped people so far so here's a real case study from a few years ago drug called erlotinib in initial clinical trials which were not targeted in a patient specific way they noticed that the drug worked really well in some random ten percent of the people they didn't know why it didn't seem to do anything to the other ninety percent so you know just looking at this from a maybe 20 year old like oncology company point of view this would be a failure this would be a failed trial what they noticed was that also about that same proportion of patients have a mutation in the gene that that drug is theorized to target okay so two separate pieces of data now if you put them together and you condition one on the other you'll actually see that all of the response to this drug is in that class of patients that have the mutation so that led to the current standard of care which is that every single lung cancer patient that goes to a hospital gets tested for that specific gene to see if they have that mutation and if they do they get this drug and it helps them out quite a bit so here's something that just happened immune therapy this is like the new hot field in cancer therapy right now I mentioned the immune system before how that's one of the many things that has to go wrong in a patient's tumor before you know before it's actually a successful tumor and the results are really interesting and they look a lot like the lot like the results for Latin have did a few slides ago it works really well in about seventeen percent of the patients in fact I believe in a good percent of that seventeen percent the patients just completely stopped and for the duration of the trial they haven't like continued to develop more tumors which is pretty remarkable so this is a pretty exciting one but it's just seventeen percent you know why is that why is it just that seventeen percent unfortunately we don't know current stratification methods have been unsuccessful so that brings me to the next topic which is pathways so how do you actually model a cell you know most cellular things aren't simple single gene stories you've got these things called pathways so here's a really simple example of a pathway on the surface of the cell you've got some gene called rtk when there's a lot of that gene it can signal this gene called pik3ca when there's a lot of that gene it can signal this one called a kt when there's a lot of that it can signal cell survival and proliferation so what happens in tumors well in a lot of tumors you actually see a mutation in this gene called pik3ca it's what I call a gain-of-function mutation a lot of people call it that not just me gain-of-function mutation it activates this interaction right here and any other interaction that I'm not showing in this graph and as you could imagine that has downstream effects okay so now it's not just one gene interacting in isolation it's now a whole cluster of genes it's a set of genes acting in a network interacting with each other talking to each other this is more of what tumors actually are right this is what's happening inside of your cells so how do you identify these pathways there are a couple methods there are a bunch of different ways you could go you could just look at the RNA levels you could try to do that without any idea of what the pathways are I'll show you what that looks like in a bit you could also try to take advantage computationally of known pathways structures that's actually a little tricky because often pathways structures aren't complete there's a lot of missing data and also in certain cell pathways are actually different than they are in other cells like different genes will talk to different genes it's kind of confusing so their drawbacks to that approach even though it sounds a lot better there are also things you can do with DNA just by itself and their methods out there now that can integrate both DNA and RNA at the same time and try to model the cell more completely but again for the same reasons there it's not clear which one of these methods is going to be the most productive and chances are it's going to differ in different cases so first off I want to talk about RNA without pathways this is surprisingly useful kinds of analysis right here this is how we are currently able to predict what kind of what kind of cancer someone has when they come in about four percent of all cancer patients the doctors aren't able to figure out where the cancer originated which can actually change what kind of care the patient gets so it's pretty useful information different cancers use different genes right so different cancers will use a different set of genes to be active and what you can do is through a feature selection gene selection this is showing differential expression where red things are way higher in the tumor than normal three standard deviations for the most read or way lower in the tumor the normal negative three standard deviations for the most blue and what you can see is that different tumor types very consistently tend to use different sets of genes so if you're really careful with your machine learning and you avoid overfitting because they're 20,000 genes and maybe only about 200 patients or so so what is that n greater than P or P grid of the net I figure out which direction it is but it's it's the bad one um so yeah but it is very useful if you're careful you can also take advantage of the pathway diagram itself and when you do this you're typically looking for consistency with the pathway structure so for example just showing a subset of this this gene here akt1 since there's a lot of it this right here is not an arrow this is a blocking a blocking relationship so if there's a lot of this you'd expect there to be not very much of this so you can do these kinds of consistency checks with different sets of different sets of genes and their known relationships of their assumed relationships and you can fit and you can score based on that and then choose which pathways are being perturbed the most based on these kinds of consistency relationships so I talked about DNA only methods earlier the problem here is that pathways have a lot of redundancy so this is the pathway I showed you earlier now imagine you had you know I showed you this specific interaction here where you had a gain-of-function mutation in pik3ca that has the downstream effect of cell survival and proliferation and it works like that you can actually have a loss of function in this gene with the exact same downstream effects works just as well for the tumor so why should the tumor really care similarly you could also knock out that gene same downstream effects why should the tumor care a lot of times the tumor doesn't care and it just gets one of those randomly so if you're doing a statistical test and you're only looking at one gene at a time you have much less power because all of your signal is just randomly split up between all of the redundant events right so if you take the knowledge of the pathway diagram and you take the individual events that occur and you spread the influence of those events through the pathway diagram and then you can use that as your input to two standard clustering methods or other kinds of machine learning methods so there's a lot you can do with pathways just based on DNA by itself so i mentioned integrated methods so this is this took me a little while to understand but this is pretty cool stuff so what this is showing what i want to show you right now is that there is signal there there signaled between DNA and RNA when something happens in DNA and you have a known relationship between one and another gene gene be the second gene in that known relationship the downstream gene if you expect that gene to be say turned on by gene a you can look at gene a being more abundant in the DNA and you can observe gene be being more abundant in the RNA right so if you think about that for a second that validates the idea that a the copy number is actually having an effect on the expression of gene a and it validates the it validates the interaction between gene a and G and B with a correlation between the two expression levels and similarly in blue here if you just look at the expression of a and the expression of B you also have higher than expected correlation out of the two so these pathways do store some real information and you can actually see it it's tough though I mean the method that does this there's a whole company built up around it it's not a trivial analysis you have to know quite a bit about you know about the different kinds of interactions and how they work together and you have to model the thing properly and i think the one of the companies i know of that does this they use a bayesian belief network to do it it's doable but tough we should put it in spark all right so what do you do now I don't know maybe some of you are interested in pursuing this more maybe not one site that I really like is called Rosalind info they have a lot of different just little toy problems along with explanations for what those toy problems are you can go through that do some code just kind of get a feeling for what the data looks like and what it means I think it's pretty useful i mean it's might be a bigger jump between that and actual data analysis then then some things but they do actually have some really you know some more complex algorithms in there so it's not totally trivial stuff Coursera we mentioned I I didn't actually know that that that one guy was teaching a class that's really cool i'm probably going to take it these classes are useful this is our continuing education so all of us in genomics pretty much do this there's one called this bioinformatics algorithms com there's a book there are some Coursera classes that's the group at UC San Diego Pavel pezzner I believe is the guy's name he's like really really good in the field there's this group at Harvard that has a whole bunch of things including a class on variant discovery and genotyping which I had a hard time finding classes on so so that's that's a cool example full disclosure this isn't my background this isn't how I learned I I was like midway through PhD and in this field and then I dropped out to join this company but anyways I have done these kind of classes to learn more stuff about things and obviously Frank's website BD genomics org avocado pick me up issues I've done a little bit there it's totally doable and Frank is very helpful so definitely do that as well and of course you can come work for driver group with me we're hiring DevOps sis admins software engineers data scientists data warehousing people actually robotics automation specialists if you are or know anyone who knows that kind of stuff that would be really cool too so yeah thank questions any questions yes graph uh-huh no it's their graphs ya know so the question is a pathways appear to be graphs rather than just graphical that is correct pathways are graphs and are we using graphics the answer is no these are you know these are methods that we're just now starting to break into and I believe no one else has done it yet in graphics so there you go you want to do something why don't you take some like cool method and put it in graphics will use it yes ah yes so the question is what are we doing for robotics yeah actually it's it's prepping the DNA and running it through the entire like what we would call the wet lab process so basically like all of the you know physical handling of the sample up until the sequencing stage yeah it's actually really important it's not just like for the coolness in the scale it's also really important for consistency robots are more consistent than humans yeah who'd a thunk all right anyone else yes so we do both I mean when you so okay the question was are we dealing with with primary sequence files or diffs the answer is that you know we have our pipeline starts with the output of the sequencer which is just the primary data so you know that's on the order of 50 gigabytes for each person in an image ish format called bcl and then that gets converted to fast queue which is the format i showed earlier and then and then we do alignment and processing and all the stuff that Frank was talking about and yeah we go through the whole a to z on that yes ah what's it like working with driver group tell us a typical day ok so you get into the office you look at your super long list of things we don't drink beer in the morning look at a super long list of like things that we still have to do because like we're doing so much stuff there are so many methods out there that just they're not implemented yet you know we've we're dealing right now is like how to take our lab what we've built here in San Francisco and expand it out to China you know we're it's a lot of stuff but you know we've got a really fun office everyone's really smart we like working together we like hanging out so there's a lot of that as well as we've actually had cancer patients come by so it's it's a very interesting mix of like you know serious oh my God we're actually helping save lives with we're having a really good time and doing data analysis and building these giant tools so I like it yes you can try hmm so the question is have we tried hold genome pcr freeze in other word you mentioned is there any benefit to that personally right now I'd rather go in the other direction for some technical reasons I'd rather get even more focused on a handful of genes that we know are more important just because you never get a pure sample okay so if you're doing a research project for example your samples probably a hundred percent tumor cells right so every single read that you sequence is coming from that tumor now in the real world what you get is some like crappy biopsy that's not taken very well and most of its surrounding normal tissue not a lot of its tumor it's highly degraded you got to go deep on that to really see stuff you've got to resample it a lot sometimes you know we've had samples come in that are like maybe five percent of the sample is tumor and the rest of it's just surrounding normal tissue so for that you really have to sequence deep so if we were to sequence that deep on whole genome you know I'm thinking like maybe a hundred X or so now we're talking not like 50 gigabytes per person it's maybe a few hundred gigabytes per person so maybe we should hire you to help us do that and also someone really rich to help us store that yes day research papers to read there's some really knowledgeable conversations like how do they happen they in front of white board in the hallway yeah well I mean you know we do a lot of the same stuff that a lot of tech companies do with Trello and slack and all that kind of stuff we also have white boards and we do a lot of just standing in front of the whiteboard and drawing stuff out and you know thinking about algorithms and how to normalize between different tumor purities between a you know tumor one and tumor to and after you've done that normalization how to correct the data for that so that you can see if a mutation is actually increasing or decreasing and abundance E and not just changing in purity between the two samples I mean there's a lot of fun stuff that we think about and wait Borden yeah yes mm-hmm yeah mm-hmm yeah it's a great question so the question was about whether or not we're doing what's called single cell sequencing from patient data and because that seems like that would solve a lot of our problems sort of at the hardware level be like pre analytic how much do you know about how hospitals work in the US and how difficult it is to to convince the pathology department to actually give you a tumor sample I mean it's it's tough for next-gen sequencing data like we're doing we're lucky if we get the drugs from a sample and that is just like you know deep entrenched bureaucracy and it is really hard so earlier I mentioned that we are opening up a lab in China yep China does not have that bureaucracy they are happy to do the like most up-to-date fanciest newest kinds of algorithmic analysis we're actually working with with the hospital over there to put a genome sequencer right next to the o.r so it's like pretty amazing stuff and you know we're going to be the first people seeing that kind of data from actual patients so yeah come work for us yes that's a good question so the question was what's next after lung cancer are we looking at other kinds of cancer it depends on the time frame you're talking about so if you're talking about someday yes we want to do every kind of cancer the problem is if you spread your focus too thin you're not going to collect enough data in any one tumor type fast enough to make a clinically actionable decision about what drugs you should go after so our company's business model is actually to figure out where there are unmet needs in patient populations with particular drugs so what we want to do is figure out for example that you know a certain percentage of lung cancer patients have certain kinds of mutations and we think they would do really well these particular patients we think they would do really well if they receive this specific kind of drug so what we want to do is kind of bring bring all that together so we're actually going to be running some early stage clinical trials ourselves and hopefully figuring out drugs that work and then sell those drugs back to the pharma companies so that's the goal