Devreal

Text By the Bay 2015: John St. John, A High Level Overview of Genomics in Personalized Medicine

Text By the Bay 2015: John St. John, A High Level Overview of Genomics in Personalized Medicine

Recording: Text By the Bay 2015: John St. John, A High Level Overview of Genomics in Personalized Medicine

so my name is John I work for a company called driver group um right now we're in the process of working with hospitals starting with UCSF to process their late stage lung cancer patients uh process their tumor samples figure out what it is that makes their tumor tick figure out what it is that's driving their tumor and causing their cancer and then provide personalized therapies for them and then when we identify situations where there isn't a good personalized therapy for a specific group of cancer patients identify the best drugs possible and actually run early stage clinical trials so we're trying to cover uh quite a bit of the uh of the typical you know going from basically early research all the way through a drug development space um but today what I'm going to tell you guys about um is first off a highlevel overview of cancer what causes it what um you know the molecular basis of it and what that means and then I'm going to talk about how we actually turn tumors into text and what that text looks like and what that text means and then um the different types of information that we extract from this text and then finally I'll get into um a couple examples of how the field has turned this kind of text Data into actionable clinical information so first off cancer is a disease of your normal cells so it's a problem when one of the normal cells in your body like a lung cell for example starts to uh divide uncontrollably and not act like a lung cell anymore right so it's it's a really interesting problem in that it's it's actually your own body your own part of your own body that's that's just misbehaving and going uncontrollable um luckily a lot of different things have to go wrong in a Cell before it'll actually turn into a tumor so um you know there are lots of different safeguards for example your immune system is sitting there just waiting for uh for a bunch of different mutations to pop up in one of your cells and when it identifies a cell that has a lot of mutations it'll attack that cell and kill it that's what happens normally um tumors figure out how to get around that process um similarly existing cell death so there are other mechanisms within a cell that cause a cell to commit suicide when uh when there are too many mutations within that cell and it's just too messed up internally um so tumors figure out how to how to evade that so those are just two of the examples of of things that have to go wrong and within any one of these there are a lot of different ways that that specific thing can get broken so it's a pretty comp there are a lot of different things that you could be looking for it's not a very simple world world so just to reiterate uh it's it's not going to be one mutation generally speaking that goes wrong within a tumor it's going to be an accumulation of mutations that slowly take a cell from a normal state where it's behaving properly um all the way through uh potentially having uncontrolled growth and just to clarify when I say a mutation what I mean is it's it's a change in your DNA in your genome so your genome is you know what it's like the blueprint for every cell in your body what makes you U so for example you might have a couple extra um letters a couple extra acts and G's inserted into a portion of your genome and if that causes uh an important Gene to break that's not good okay so the DNA of your genome that's that's huge that's three billion base pairs the parts of that that actually get turned into genes is about 1% um we call that the exome um so basically there's this up here there's this standard flow between um between DNA of of the genome um you can focus in on just genes that get turned into RNA which gets turned into protein proteins are the things that actually do stuff inside of your body so the proteins are you know they carry signals between your cells they make up the structure of your cells uh they make the cells do different things so um when you have a cancer what's actually happening is you're getting mutations in your proteins or the part of the genome that regulates how much of those proteins are present and that causes things like you know for example here perhaps this mutation in gene C Gene C is a little bit messed up maybe that caused a gene a to turn off when it should have been on and it might have caused Gene B to turn on when it should have been off so you can get these kind of signaling Cascades ades that happen inside of tumors these kind of like networks of cross talk between different uh different genes that get mutated in different ways and signaling other genes that can cause a cell a normal cell to start misbehaving so DNA is huge again um if you took the DNA in one cell even though it's only about 10 atoms wide a strand of DNA only about 10 atoms wide but if you unraveled all the DNA in just one cell of your body and again cells are so small you can barely you can't even see them unless you're looking at it under a microscope um if you unravel that it would be 1.5 M long so we're talking about a huge amount of of information just crammed into a very small package so how do we turn this into text um well I told you we do a uh a normal versus tumor comparison so we actually sequence both we're looking for um for differences between a person's normal genome and their tumor genome uh we're not interested in the things that just make a person unique as a person we're interested in the things that make their tumor unique as a tumor so uh we extract uh specifically the exome which again is the part of the genome the 1% of the genome that actually gets turned into proteins and we extract the RNA and uh we throw this into a sequencer so this is about 55 million acgs and T's that we um that we extract for one copy of the exome but we have to sequence it many many times over so that we're really confident in our observations and for a couple other reasons I'll get to in a minute so how much data is this well if you imagine Warren peace super huge book about 1,400 pages of text 3 million letters um we sequence about 6,000 copies worth of that uh for one uh person so this is this is a lot of data to go through and because of a uh because of a current um I guess like a current limitation in technology that I don't really see getting over any time soon um we have to do the equivalent of taking those bookshelves filled with books and throw them into a paper shredder more or less so what we get is this just big pile of data without any context uh which is a lot of fun to go through so what are we looking for in this big pile of data well there are about 100 to a thousand new mutations that show up in a person's tumor sample right um of a uh in the DNA itself you can get slight shifts in how much of the DNA is present so you can get like extra chromosomes or parts of the chromosomes that are lost so we have to identify those situations those can be really interesting too um and then also the RNA which is a surrogate for how much protein is present could be there at different amounts so we have to pull all that out of this out of this text data so what does what does this look like let's zoom into one of these little little fragments in this big pile here uh it looks like this you've got a little unique identifier um that just tells you what that sequence is and where it existed in the imaging software that found it U the DNA sequence the A's C's G's and T's and then uh individual confidence scores on each one of those letters uh because um the sequencer itself has a little bit of knowledge about how how confident it is in every single base pair that it reads so you can use that when you're trying to figure out what the data means and how much you believe things so this is the first kind of thing that we look for um mutations so what I'm showing here uh one of those little gray bars is actually a um a sequencing read one of these that I just showed you aligned to the genome so we give it context we figure out where it most likely goes it's kind of like a Google search for example where you just figure out like what it matches up with the most um so you do that for the normal sample and the tumor sample against a standard human reference and then you can compare the two at every uh position in the alignment and you can identify situations like this where a certain set of the uh sequencing reads A Certain set of the DNA in the tumor has a a difference that isn't there in the normal sample so for example this right here is showing you a chunk of DNA that was actually deleted in this person's tumor and this specific one is uh pretty important um if you have that specific mutation in your tumor you're eligible for a drug called a lot nib which uh extends your life pretty substantially so um these things are really important and we have uh we have algorithms that identify them um I mentioned copy number alterations this the second thing in DNA that we look for so this is when for example you have in one of your chromos rosom a chunk of that chromosome is just duplicated it's there twice instead of once so how do we find that well actually we look back at that alignment between the normal uh and the tumor and in bins we just count how many uh reads are there in the tumor versus the normal and you get a ratio for that bin and um what you can do is you can plot out these ratios across a chromosome and you can identify stretches where uh whole bunch of sequential bins have a shifted mean of their uh of their ratios and this is a pretty strong indicator that this person has a um has an amplification an extra copy of DNA at that position so um there are a lot of drugs that you can actually recommend just based on those kind of changes and finally RNA this one's tricky so instead of um instead of now showing you an alignment to a specific genome segment what I'm showing you is an alignment to a specific Gene so you would count how many reads of uh of the RNA are matching up with a certain Gene in the tumor versus the normal um and you can plot this out sort of uh you know XY plot of normal versus the cancer sample and um there's a lot of noise in this kind of data unfortunately you've got about 20,000 genes there's a lot of dispersion um a lot of different genes are just going to be like randomly overexpressed or underexpressed it's it's really complicated stuff but every once in a while you'll find really high magnitude outliers that you can you know really uh hang your hat on and uh those are the situations that we that we look for um and also highle patterns so different tumor types will actually have different sets of genes that that tumor type uses to um to basically uh to be a tumor and that might be different for from lung cancer and breast cancer for example so um we've actually used that to categorize uh tumors of Unknown Origin where the doctors had no idea where the tumor came from uh because it had already metastasized throughout the patient's body uh so these tumors were everywhere but we were able to look at the RNA in that person's tumor and compare it with this huge public resource and figure out oh this actually looks the most like a Cervical Carcinoma for example um so there are a lot of really powerful things you can do with this so now I'm going to shift gears and talk about a case study so this is uh something that happened about um oh no so um this is something that happened about 5 years ago um there was this drug I mentioned it a little while ago or lnib that um in a very small set of patients when you give it to all lung lung cancer patients in a very small set about 10% it really increased their survival so in that subset of patients you know they had really really great survival but when you averaged it across all patients it didn't actually look like a very promising drug so researchers were wondering um whether there is a way to figure out who is going to benefit from this drug versus who wouldn't um who should they just give regular chemotherapy and radiation to versus who uh they should definitely give this drug to um so what they um what they also knew was that this drug targets a specific Gene called egfr so they thought well maybe if we could figure out which patients had tumors that were dependent on the gene egfr to be tumors then maybe they could stratify the patients better so here's what egfr mutations look like in lung cancer they affect about 8% of uh of Europeans it's actually different with different um racial backgrounds interestingly um and uh in the other 92% of lung cancer patients are completely normal in this Gene so when they did the stratification when they took the just the patients that had a mutated egfr versus the ones that didn't 81% response instead of eight so this was definitely a clear way to stratify these patients so this is actually the current FDA approved test um if you have lung cancer uh you go to the doctor they'll do two tests one is for this Gene and one's for another Gene called ALK um to figure out if you're a candidate for either of these drugs versus just chemotherapy all right so what about right now what about a current problem s in the field so here's one this is actually a really exciting new area it's called immunotherapy earlier I mentioned that the immune system and evading the immune system was one thing that you had to do uh to basically be a successful tumor so um drugs that Target the immune system are actually super effective in about 20% of of lung cancer patients the rest of them not so much however um when people really drilled down on the single genes that have been involved in uh that these that these drugs are targeting they haven't really found any correlates yet so if you just look at that one gene there aren't any correlates so what I'm guessing is that this is going to be a much more complicated uh text mining problem data science problem where it's going to be a more more interesting interaction between many different genes potentially that give you the signal for whether or not this tumor is going to uh respond there's some indication that maybe just uh the mutation rate how many differences there are in the person's tumor is an indicator of whether or not they'll respond to this drug but um it's still very actively researched and there's a lot of debate and no one really knows so if you want to get involved that's a cool area so anyways text mining the kind of stuff that you use to make recommendations on web pages and all that other stuff you could also make recommendations for curing cancer I'd like to think that's cooler um so where do you go from here let's say you're interested in this and you want to get involved um you could go to websites like rosalin doino they have a lot of bioinformatics resources you can learn about the field and the different algorithms and all that um you can contribute to open source projects like uh the amazing one Frank in this room is working on um go to BD genomics. org or listen to his talk next um and also potentially could come work for me at driver group so we're looking for devops people CIS admins software Engineers data scientists so come talk to me afterwards if you are curious or interested um and now open it up to questions is the hypothesis that the difference is the only thing that matters a robust intested hypothesis or is there like potentially unexplored linkages to the underlying uh genetics of the person to begin with both definitely both right because um you know I think one important clue is that uh for example in different racial groups you get different levels of egfr mutations that suggests that there are like background differences in just people's DNA that uh that correlates with whether or not they're likely to have a specific kind of lung cancer um there are also things like familial tumors that are um you know you you inherit certain mutations that make it more or less likely that you develop cancer by a certain young age or um there's braa mutations probably heard of that they do braa tests now for a lot of people to figure out whether or not um you know you're likely deel to develop a breast cancer um if you have a mutation in that Gene you're I think like 20x more likely or something even though it's still a small number yeah so um yes I was recently a presentation by the CEO M mors help and they have this approach of they using super Computing entire Human Genome and then try to look at the cancer cells and try to figure out you know there's a protein sequence of some sort that s that they can then Target for specifically rather than just blast everything using chemo or or something like that uh how is you know any what's their approach what they trying to do and how are you guys different I'm not super familiar with their approach I think the big difference there other companies like Foundation medicine that are also uh providing tests like this um I don't know how much of the genome nand Health covers um we do cover the entire exom instead of just a targeted panel of known tumor related genes because we're interested in research and discovering new things um but I think the biggest differentiator is the fact that we're also going after early stage clinical trial so um this service that we're offering actually for free to patients this isn't our business model uh these reports that we're offering to patients for free is basically our Gmail if you want to think about it that way it's how we're uh getting data that we then use internally to make better decisions about which drugs to go after yeah yes how much of your work is uh I guess like Compu for technique bound versus limited by actual real world data I'd say most of it right now is uh data limited and I think the one of the most challenging things is finding data sets that are well annotated so where you know for example that um that a person responded a certain way to a certain drug and that's part of why we're really focusing on collecting these data sets along with is Rich um responsive information as possible so how do you really do that do you focus say you go to a hospital and collect all the data or ask people around the country for them to Source it or so right now we're we're focusing on uh Partnerships with specific hospitals so we're starting with um UC San Francisco and we're working on expanding that to the entire UC system um I think we're in some talks now with some hospitals in China as well so um yeah yes have another example for the relationship text Antics and sequen um I don't have other examples right now um one of the things that's a commonality is the uh the stacks that we end up using so um for example Frank in a little bit's going to talk about uh his tool which we're working on uh incorporating into our own uh into our own pipeline which uses Spark and a couple of other um things like that but yeah I mean really it's all we turn these tumors into Text data and then at that point you know there are a lot of the same kinds of algorithms that you apply to uh to the machine learning problems um with this kind of text Data versus other kinds of text Data MH yeah yeah yeah so your short example is a single single G SLE is so if you do have and you also give a c example say that example is not due to a single liation or something so in that second case what is the general approach you will take to problem yeah so um I would say uh it's a tough question you know uh because there are a lot of people working on this right now I would start looking at a lot of different features and develop hypotheses about you know these different features and then try them out so one of the things that we're thinking about doing internally right now A lot of people are just using the raw counts of mutations but mutations can appear in tumors at different frequencies for example so um for example maybe every single cell in one of your tumors has a mutation versus maybe only 10% of your cells in your tumor have a particular mutation versus you know the other 10% have another mutation another 10% have another one and and so on so you can get different levels of what's called heterogeneity inside of your tumor so it might be that uh even tumors with a large number of mutations if they are heterogeneous versus not heterogeneous they might actually uh respond differently and you can imagine that maybe the immune system could attack a tumor cell better um and you have to think back to like what it is that the immune system is recognizing it's recognizing mutated proteins so it might be able to attack and learn how to kill a tumor better if all the tumors are showing the same set of mutated proteins versus a whole bunch of different kinds of mutated proteins right so um that's one of the things that we're that we're working on right now are uh pretty Advanced normalization methods that can answer those kinds of questions whether or not mutations are mostly at high frequency versus low frequency all right thank you