Devreal

data.bythebay.io: Atul Butte, Translating a Trillion Pointo Diseases of Data into Therapies

data.bythebay.io: Atul Butte, Translating a Trillion Pointo Diseases of Data into Therapies

Recording: data.bythebay.io: Atul Butte, Translating a Trillion Pointo Diseases of Data into Therapies

so F first of all there's more of a bio online if you want uh I'm a medical doctor um and all the slides are online and there's hours of video on YouTube so don't feel like you have to like take screenshots and stuff everything's uh publicly available uh but first i'm a medical doctor usually presenting to audiences I got to talk about my conflicts of interest so I just have a few suffice to say I've started a bunch of companies I consult for a bunch of companies you probably don't want to believe another word I say over the next 20 minutes wouldn't blame you I'm in Professor mode already with my stick here by the way um but uh I'm most proud of the bottom right those are all the companies started by my students more than half my graduate students start companies now uh they do this even if they become academicians you know professors they should still start a company anyway uh and they do this with the best platform in the world and you know what that is that's data it's often Big Data it's often open Big Data so I'm going to talk a lot about open data here open Health Data so I'm going to show you how they do it I'm going to show you how I do it maybe I'm going to convince you this is the most amazing time to be in Health Data Sciences right now it's just unbelievable time we're in this uh amazing mode now of trying to get to Precision medicine so President Obama himself me Mees Precision medicine the State of Union Address a year ago and what that really means is we're going to be able to we're entering this era where we're customizing the healthare we deliver to a person based on the measurements we're making on that person right it could be molecular measurements DNA could be your Fitbit could be wearables but we also have to take in account all the measurements we're making on everyone else right it's not just you you you but we got to learn from everyone else that we've been taking care of and then apply it to you so we got to record what we're doing to all the all the patients out there it's public health but just bring it uh into you and you can see the three reasons why the government thinks it's the right time it's sequencing of the genome which is data Rich biomedical analysis data rich and new techniques for using dollar J and they literally show a picture of the cloud I mean the three and only three reasons the federal government wants to do this is because we're in this data era right I mean you can't get better than the president of the United States like pushing for this right so we're in this data delusion healthcare because of amazing gadgets like this now I could have showed anything I could have showed you Mo's law plots could have showed you a sequencer I love showing these because this is how I got my start in the field with these little kind of nifty gadgets uh you put any sample on they you got to process it first but you put a cancer sample diabetes sample a mouse model if you're into that and gets a read out of every Gene in the genome like how much of each recipe is being made a lot of this one not so much of that one now when I got my start in 1998 1999 these were Priceless we had like 60 of them 60 kind of cancers we were analyzing that and now they're so cheap I carry one in every suit just for talks okay uh it's also the thing is you know big data is kind of ethereal so I like to show Big Data comes from small packages just like this right in fact now everyone runs these chips okay it's not even a big deal in fact when you get tired of measuring them one by one you measure them in the 96 well plate format and that's so cheap I have that in my bag this one's kind of falling apart a little bit but you pretend you see 96 squares there right and then when you get tired of measuring on 96 at a time we now have the 384 well plate format and that's so cheap I can put that in my bag I cannot more clearly illustrate exponential growth in biomedical data than to show you these plates right I started with one by one and now we don't even use these anymore uh because we moved on to sequencers in fact this company doesn't even exist anymore they just got acquired AP metrics like 3 months ago that's how fast this field Moves In fact so many people use these chips that everyone started write a scientific paper a publication right here's a list of genes that change in that disease and this disease and heart failure and cancer and they all started to write their papers submit them to journals and what did the scientific journals say they said hold on a second we're getting too many of these papers no more papers with these chips in them no more okay unless you put this data into an international repository where the reviewer reviewers can double check to math it's probably important to do that we're getting to patience now and maybe others can use these samples for something else so this precedent in science goes back almost 38 years we have something called gen Bank you may never have used genbank you might have heard of genbank it came with the first DNA sequencing papers you know there's a figure one has AG GTC these letters of the DNA nobody wants to type that in again right so they're sending tapes to each other right this is even older than 8 in discs if you remember those which are older than 5 in discs 3 and 1 half in discs CD ROMs DVDs and of course the worldwide web so for 38 years scientist been sharing raw data and the precedent is there if you get enough scientists to use the technology you got to share the data eventually fast forward to August of 2012 uh this article came out of nature one of the magazines what happened in August of 2012 we hit a million samples publicly available a million cancers a million diabetes a million biopsies someone already ran the CH got the sample someone already processed it someone already digitized it data freely available on those samples now this is deidentified data you cannot figure out who these people are but you can get an idea of what that sample is do what that tissue is what that disease is doing you can see the growth C from zero to a million in fact if you look really closely it looks sad because it looks like it's slowing down there at the end and then you realize they only counted half the year that was August of 2012 okay still doubling four years we're just a couple thousand short now of 2 million publicly available probably by the end of this month or next month we'll hit 2 million samples publicly available so what does that mean right any of scientists at UCSF or Stanford or anyone any good scientist can go now and start of starting at the bench at the wet bench with a pipetter they can start looking for data right on cancer diabetes forget our scientists even a high school kid now that needs to do a science fair project she can go to these websites for for example if she needs to do a science for a project on breast cancer type in breast cancer hit search and find and download nearly 70,000 samples of breast cancer for her science fair project about as easily as she can find a song on iTunes today okay 7,000 samples of breast cancer already ready the what's so magical about that number 70,000 if you haven't figured it out that's more samples of breast cancer in the repository than any one breast cancer research researcher will ever have in their lab because every one of them eventually has to send data into the middle and whoever figures out the middle now has more data than actually any scientist in the field what and if it's not breast cancer you like maybe it's colon cancer prostate cancer they're all there waiting for you because can I tell you these scientists when they get their grant money and they go make their measurements and they go write their papers what do they do next yeah they submit the data and then they go get another Grant to go get more money to go get more samples done nobody in this field in generally knows how how to use this data and just sits there doubling and nobody's using this data so I'll give you two anecdotes of what we did with this kind of data the first is a idea first big grant that I got when I was still at Stanford was go get every human disease studied by these chips that we public data for cancer diabetes not just cancer what kind of cancer we started to collect experiments of disease and healthy disease and healthy and all that data is in the repositories we use text mining to figure out which disease was which we human eyeball it to make sure we got the right diseases in the right place and you know I'm not even going to show you genes and chemicals and molecules I'll show you these pretty little boxes the boxes each one is a gene or recipe in that tissue and the colors maybe mean more of this or less of that and we uh we calculate these signatur so signatures means what's changing in that disease from healthy to normal right what's the difference there so what are we going to do with this well if we got an idea of what's going on in the disease at this level maybe we can use that to make a blood test a diagnostic what's a diagnostic it helps us tell as doctors what kind of disease a patient has right what do they have in front you might have Imaging diagnosis like an x-ray or blood tests and so we wanted to use this data to help us make blood tests so we did a lot of these but the one I'll spend a little bit of time talking about is this this disease called preclampsia any of you have heard of this disease preclampsia show of hands yeah so it's usually about half to third to half the audience it's a half that actually watches down to Abby because one of the characters dies of this disease spoiler alert if you haven't watched it actually they dies eclampsia right so you can be pregnant you have high blood pressure you start seizing that's kind of a nasty thing to happen preclampsia happens even before that and still how husbands often lose wives around the world okay it's a nasty disease and believe it or not our diagnostic for this is the most non-specific thing in the world we just look to see if if there's any protein in the urine it's just so crude and still this disease just keeps happening and a lot of people might have friends or family have had this happen there's no good diagnostic for this so we want to come up with diagnostic so there three names at the bottom there Linda L is a grad student Matt did uh Bruce did the protein work you'll see in a second and Matt got us involved at the end here and so the idea here is let's just search the repositories for preclampsia okay so let's search here's 266 experiments here here's one of pre-term birth preeclampsia preclampsia uh prematurity how many more do you need to get started here someone's already studied this for you they've run some got the samples ran the chips deposited the data and so what's the method ology I'm not going to show you any code I'm not going to show any formulas just imagine a massive thenen diagram right I'm trying to figure out what's in common across all these experiments what are they all seeing I don't trust any one of these experiments but I trust what they get in common right that's wisdom of the crowd here I don't trust one of them I trust what they all see in common so we chase down what they all see in common and we start to do some protein work and here's an example of a new protein blood test that works so it's called hpx hemopexin the pink are the pre clamps of women the green or the non- preclampsia women and you can just by I see that it's higher just as we predicted here we found like seven or eight of these we file patents we wrote papers and what do you do next in Silicon Valley you start a company on this okay so you got the unmet need you got the data we had marchad dimes funding the early work uh we analyzed the data designed the diagnos we got one of these seed grants at Stanford called spark we have them at UC called Catalyst they give about 50,000 bucks if you have a good idea we got the test to work uh with the the samples in the repository and then uh spun a company called carmenta raised $2 million seed financing now that's an interesting story this was a oneman company Matt Cooper we launched a multicenter trial because you got to test this for now prospectively and why do I start with this one because I love this story we already sold the company it's already been acquired so we went from public data all right free data you don't even need a user ID and password to get to this data designed a diagnostic test Ed it in in the lab launched the company got seed financing sold the company 24 months start to finish inventors happy investors happy and we're going to do more of these you bet we're going to do more of these I'm giving away my secrets guys because we need so many Diagnostics in medicine you could all pick a disease and we'd never step on each other's toes okay so that's a diagnostic story from zero value public data to acquire an acquired company in 24 months second stories on drugs and pharmaceuticals the story is much worse there you know if we think we don't have diagnostic we don't don't have enough therapies Therapeutics because everyone thinks it cost a billion dollars to develop a new drug it's actually that's an underestimate it's not a billion dollars it costs way more to develop a drug you know by the way how do you get the right number of how much cost to make a drug how much did you spend divide by how many drugs did you get right simple math what did you spend divide by the number of drugs and for these top 12 companies it's between $4 billion to 12 billion per drug kind of not sustainable okay you may love the farm industry you may hate the farm industry look I consult for them all I love them still but I don't have faith that they're going to solve this problem at all okay uh it's just not even the right culture forget about picking on the farmer industry for a second even if every farm and biotech company on the planet was successful they not enough of them to develop all the drugs we need for this modern era of precision medicine so then you realize it has to come down to us to solve this problem typically in Silicon Valley like we cannot keep it going at 12 billion we got to figure out cheaper ways to develop drugs so how we're going to do this we're going to use public data that's my approach to do this and you know while we're turning our crank we started to realize if we can collect the disease and healthy on the left we can start to mine the experimental data for cells treated with and without drugs here are some cells in a pach with the drug and without a drug and they run the same kind of chips that get the same kind of measurements and we're going to put these two together the reporters for this call this match.com for drugs okay you know Match.com and what's a famous saying when you're looking for a spouse or a mate right Opposites Attract so here's my informatics methodology if I've got a disease on the left where this Gene goes up and this Gene goes down and I can find a drug that can make this one go down and this one go up maybe there's a match there it's as stupidly naive as that okay I'm just looking to reverse the profiles that's two genes imagine 20,000 je you can imagine karov smof test and all that stuff right so we turn our crank we got lots of ideas you know here's a drug there's a drug but actually where we get a lot of traction is here's a new use for an old drug this idea of drug repositioning because these are drugs that we already know the safety profile for but maybe there's something else that they're actually good for so we turn our crank we get lots of those uh and this we've done more than a dozen this is my favorite this a on Publications all the stuff is open access you can get to it here's a drug called emine that we use for depression okay it's a called a tricyclic anti-depressant we use it we we used to use it for depression we actually don't anymore we use proac these serotonin rupic Inhibitors and this is a mouse that has lung cancer you can kind of see the lung nodules there here you're not got these genes the lung the mouse gets mouse lung cancer here and on the right is the same kind of mouse with the lungs all better with the amiine now we don't use amiine anymore because it's got some side effects it makes people sleepy and uh it might set you up for an arhythmia it might set your heart rhythm a little bit off uh causing a little bit of a disturbance in your heart rth them actually neither of those two side effects sound as bad as having lung cancer 5% survival rate only 5% of people survive after this diagnosis in five years and the cancer has melted away here it's gone in this mouse now why do I love this one okay this one's amazing because we start with public data we get the prediction this drug is going to work on lung cancer we test cell lines in a pachy dish not going to show you that we got this lung picture 15 months after the computer prediction we got the ethics board approved we launched a clinical trial on on this one humans are dosed with this drug now this is a computationally driven clinical trial 20 patients recruit on the study total cost of this trial 50,000 bucks okay not even a million not even 100,000 bucks 50,000 bucks gets us this trial going we got to do way more of these okay and not spend a million or billion to do it we got to get way more of these to happen no pharmaceutical biotech is targeting small cell lung cancer today this is zero none of them are there's just too many other diseases they're working on and what do you do you got to start a company so this one is new Med we' raised more than 5 million bucks on this one and I love this story so of course my wife runs this one so I love this one the best out of all my companies got love this one but you know this little team of seven people uh over in paloalto kind of like in a garage I call these garage biotechs because we love garages in silon Valley they're developing uh psoriasis drugs for Allergan so until last month Allergan was going to merge with fiser and be the largest pharmaceutical company in the world now this team of seven people full-time is developing drugs for Aller the Behemoth Aller actually it just takes seven people to actually do this with you know enabled with a shitload of data right that's the idea here that's new Med and they're hiring data scientists right now too all right where are we going next in the last four minutes what's the next big open data I can't spend a lot of time on this but it's clinical trials data the whole world is heading towards this next big thing I'm skating to where the puck is going it's here because FDA is going to make the pharmaceutical company share raw clinical trials data it's going to be an amazing Treasure Trove you're going to look at failed trials successful trials I'm showing a URL here because I run this one for NIH we give out we're the only website where we give out more than a 100 raw clinical trials data every patient every drug every arm every encounter deidentified you can't figure out who these people are but you get the idea of what a trials data set looks like learn now because I see where it's coming half my research group is just on clinical trials data now and to could sign up at this for free account here and learn how to do it now I'm lucky because I got recruit the UCSF from Stanford a year ago with building this institute for computational Health Sciences we're kind of sluming it out in this brand new building until they build our brand new building right next to where Illumina is the sequencing company so you couldn't even buy a better location than that but of course the warrior Stadium will be right next door so maybe we'll even get some decent restaurants down in Mission Bay because right now it's kind of a restaurant desert if you've ever driven by there and that's kind of cool but why did I move from Stanford I love Stanford but why did I move to University of California is because UCSF is one of five University of California medical schools right there's UCSF and UCLA those two are behemoths got our Irvine Davis in San Diego and they actually put me in charge of organizing all 14 million patients that get any care in the entire University of California system that's 4% of the US population gets some care in University of California system I'm putting all that raw electronic health record data together in one database by the way we're already done I wanted this done before my first year so it's living in fisma compliance over in San Diego supercomputer Center on one big dbox running SQL Server pains me it's running SQL server but it had to be that and we have about for example uh nearly a billion lab test measurements on these 14 million patients that's how much data we have we imagine all sorts of amazing uses of this data we're going to enable not just us but enable the entrepreneurs enable the the activists the social researchers the cancer genetics researchers the chief medical officers the app designers we have a mandate in the state to really enable everyone not just ourselves here with this day we got to slice and dice and de identify it of course keep it safe we got to keep it private I'm going to build maps with it kind of like the last speaker I love building maps of death and disease that's kind of morbid but this what I mean you know like what do we predict is going to happen next with our patients so here you see a map like alcohol mental disorders that's alcoholism and then a year later some come in with liver curosis and then a year later a bunch come in with liver abses and then the squares mean they died of that disease so you can see some people go straight some people make a detour they get diseased they get diseased they get death so I'm kind of figuring out death and disease here's a harder one heart attack a lot of patients with heart attack die some get heart failure lung disease and some die subsis which is a bacterial disease I never knew that happened after 3 years of having a heart attack and what am I trying to do this is work by Hannah and Jay you can see their names there and our brand new building in Mission Bay we're going to have a wall of monitors there and we're going to show our maps of death and disease but better than just showing the maps I'm going to visualize all 14 million patients in the University of California system here this is a prototype of every single patient going from disease to disease to disease to death and we're going to track all 14 million of our patients we're going to predict what's going to happen in the next 90 days we're going to predict what's going to happen the next year and we're going to do something about it to prevent that mortality bidity and that's what we're building in University California system I have to thank a lot of people who help me with the work the data warehouse team a lot of collaborators and a lot of support from UCSF and IH to do all this work so thank you very [Applause] much w