Devreal

The Ingenuity Biomedical Knowledge-Base: Advantages of Modeling Knowledge in an Ontology

Event: Ontologies

Text By the Bay 2015: Jeff Lerman, Biomedical Knowledge-Base: Modeling Knowledge in an Ontology

Recording: Text By the Bay 2015: Jeff Lerman, Biomedical Knowledge-Base: Modeling Knowledge in an Ontology

so I'm deff lurman I'm um uh part of kyogen now which is a life sciences company um I uh will be talking about uh uh the ontology that we use for uh organizing biomedical information that we curate from our variety of sources and I'll talk about uh why we do that and um and how we do that uh I actually started uh when K well when I started I wasn't working for kayen I was working for Ingenuity systems U which is a bioinformatics firm and um that was acquired a couple years ago by kyogen which uh has broader scope so I I might be flipping back and forth been talking about Ingenuity in kyogen but um it's all the it's all the in some sense the same company uh all right so um so kyogen um has a biomedical ontology uh and I'm going to sort of run through for you uh what we are able to do with that ontology um where we use it uh the requirements that we have in order for it to be useful for us uh what it looks like um how we build it and by building it I mean um um structurally and in terms of the content I'm going to be somewhat light on the uh software implementation details I'll I'll take questions on that um at the end but uh this is going to be more of a uh informatics uh uh set of questions rather than uh the software implementation per se um and how we do quality control on the content which for us is uh is very important um so I'll start with some some examples of the kinds of questions that we can answer uh with the with the ontology that we have and so uh one one might be I want to get a list of all lung diseases that aren't cancers so um we get about uh 20 plus results from a question like that um and they range from the we can read these right good okay they range from uh things that might more or less consider to be obvious bronchitis um and uh and by the way for each of our our uh classes or objects if you will uh we have descriptions that are internal for us but that help us um sort of reflect the meaning or know what what we're talking about and what the meaning of these uh concepts are so um inflammation of part of the lung um infectious lung disease this is kind of a high level classifier uh and we get into um some less obvious things that you might not think of immediately when asked the initial questions so fungal pneumonia um pilosis which is a particular kind of essentially poisoning um that affects the lung and a whole range of other things that get sort of less and less obvious and uh these are you know easy to to gather um in the anology so um a more sophisticated uh question or at least a more complicated question um be to list the active clinical trials uh So currently active clinical trials for drugs that treat cancer and patients that have mutations in BFF so BFF is a is a gene um and there might be particular drugs it turns out there are particular drugs uh that are in testing and some some level of testing to treat cancer in patients that have specific that are indicated for patients that have mutations specifically in that Gene so uh we get uh we run this query recently we got 175 results over a collection of five unique drugs and I've just listed some examples here um so in order to do that we have to have models for a number of different concepts within this sentence and I'm not going to talk about uh sort of the mapping of a sentence like this because I know in a you know a in a conference like this there's a lot of uh interest in mapping sort of natural language obviously to uh to structure language and I'm not going to focus on that particular mapping for a sentence like this we would pose a question like this in terms of a query language uh directly um so nonetheless we want to know about uh the concept of drugs um we want to know about uh which of those drugs treat cancer so in order to know that we have to know about uh we have to have a full list of drug approvals uh from some agency or agencies uh and specifically within those drug approvals we need to know what the Dr drugs are approved for um we want to know some details of the in of those uh drug indications or what drugs are approved for uh in order to cover this uh restriction patients that have this particular kind of mutation uh and specifically for those indications it's not enough just for just for us to have um descriptions we want the gene or the mutation to be broken out uh in a structured way uh and then finally we need to know uh what does it mean for there to be an active clinical trial so we need to know um we have to have a full list of clinical trials we need to know which ones are active we need to know which ones uh are for which drugs uh so that's a that's a bunch of different areas of knowledge that we have to be able to tie together so so and we can as you can see we can do that with the with the structure that we have and with the collection of knowledge that we have so um Ingenuity now as I said kyogen um makes s where to contextualize molecular biology and genomics data set so the general idea uh the sort of overarching principle is that we're aligning customer data to the world of relevant biological knowledge uh in order to give insight into the meaning of uh what would otherwise be raw and and perhaps uninformative data um so we're facilitating insight into existing knowledge and the inference is powered using uh multiple individual what you might call facts what we prefer to refer to is uh findings or assertions um we're not making a judgment as to truth but we are making a judgment as to um what things are said by whom um and I have a a couple of examples of uh main software products um that still bear the Ingenuity name um one of them is um called engineering pathway analysis it's been around for quite a while now and um that's looking at the interrelationships of U or looking at biochemical U Pathways and how they interact and how uh genes play a role in those Pathways um speaking at a high level uh and a newer one is uh Ingenuity variant analysis and the focus there is to look at uh genetic variations or what you might what people would would refer to as mutations perhaps um and uh specifically to be able to filter lists of uh genetic variance in particular cases in particular patients to isolate um what might be likely causal variant for a particular problem uh that that patient is presenting with so in order to be able to do that to sort of serve as the underpinning for these this kind of software we have a number of requirements so one is um the knowledge has to be organized in a way that's accurate it has to be highly accurate we have a very low tolerance for um for outright um mistakes um and we need to be able to confirm obviously that accuracy so we need to have a framework that's uh highly testable uh things need to be normalized so uh we'll be um not surprising to people here we um so one way that we mean that is that we want to be able to map uh synonyms to a concept so we're modeling Concepts but we want to have um all the tags uh or many uh um of the terms that are used to refer to those Concepts so we can recognize uh across different data sources of course that uh a term is referring to uh the same thing um we also at a somewhat higher level need to be able to map different syntactic structures to the same meaning so if someone's referring to inflammation of meninges or inflammation of um minks uh we need or and there's more um we we need to be able to know that that means the same thing and there are extensions past that uh where we also have to uh even when the uh semantic structures are different we still need to be able to recognize that we're talking about the same thing uh so for example there may be a process uh a biological process that results in um some property of a cell being increased and so one person might say um I observed in in scientific literature I observed an increase in that process I observed an increase in uh program cell death or um in term is apoptosis uh another person might say meaning the exact same thing uh I observed an increase in apoptotic cells so these are it just means cells that are undergoing this process different syntactic different semantic structures uh we need to know that that's those are equivalent and so we need to be able to to tie those things together um needs to be sophisticated enough to be useful um so uh we need to be able to do inference across different kinds of statements uh statements with different kinds of structures uh simple example um a causes B to bind a c um the the result of that so BC binding uh increases some other process D so we uh need to be able to infer or want to be able to infer um down a chain like that so A increases d uh and so that requires us to have some kind of a unified model for statements uh even when there are a you know including a variety of kinds of statements so in addition we want to be able we want something that is efficiently extensible uh we need to be able to add new relationship types uh fairly easily so I've listed some examples of the of kind the kind of relations that we um um that we record here um physical part of um implicit locations and I'll talk more about that in a moment um whether something has a particular mutation so um within the ontology we have um a variety of different kinds of Concepts um we have physical things um at variety of scales we have to be able to talk about what we call temporal things or processes so uh diseases or other uh pathological processes as well as um just general biological processes uh and um other kinds of things as well both properties which might uh modify some of those physical or temporal things uh but also um we need models for statements themselves uh for metadata how we uh know a particular thing where a particular assertion came from um as well as definitions both in the um purely text based sense for the use uh for use internally to um help us communicate internally but also um to be able to tie different kinds of assertions to each other and and uh and bind that equivalence so um in addition to to those uh kinds of Concepts we also have a list of types of relationships uh the most frequently seen kind of relationship is that is a relationship so something that is more specific than something else cancer is a disease um other examples a physical part of um finger is a part of hand Chicago is a part of Illinois um or um at a different kind of level definitional agent so um bacterial infection we want to be able to infer or or we want to know that by definition bacterial infection has the is caused by uh bacteria um and there are a variety of other kinds of uh definitional relationships uh that we can infer to um to make explicit the relationships uh the the fact that some Concepts can be described in terms of other Concepts that we know about um so the first list of things we uh we model as classes the second list we model um is things we call that are called slots um we also have um instances of um of classes and so there are specific references to um to each of these things uh and then finally we can uh talk about scalar values so uh Str are the are the most frequently seen example but also numbers booleans um and we can uh talk about sets of of items as well uh so we use a frame based ontology um and so all of the first kinds of things are recorded as um as what we'll refer to as frames and what are referred to in a frame based ontology as frames um so this goes back uh I guess the first uh um example of or the first description of a frame based ontology was um the 70s um by Marvin Minsky our ontology was uh developed um about 15 years ago um and so after there had been uh considerable discussion about this but um anyway we we've found this to be u a very effective um model for our for our use so here are some examples of how we organize things in this ontology uh so we started at a pretty high level um say physical thing um and within that we might have anatomical part and we have so the gray uh excellent um so the gray uh items here are classes um the lines between them will be slots uh and I'll color code the slots for the different um kinds of for each each different type of slot that I'll be referring to we'll get a a different color code so um these red arrows will be uh Isa um moving down to something more specific we have the hippocampus might also have the brain those are both anatomical Parts um they also um as you might know have uh another kind of relationship so the hippocampus is also a part of the brain uh so we would we would uh denote that as a physical part of relationship um and then there are temporal things so disease uh is a tempor is a temporal thing or is a process cancer is a kind of disease brain cancer is a kind of cancer um and we can uh denote relationships across these kinds of highlevel um distinctions or high high level areas um as appropriate so U brain cancer obviously has a relationship to brain um we would call that an implicit location so uh in that way we can gather all of the processes for example that uh would be anticipated to occur uh uh or be limited to occur I should say in a particular place um and um and know for example that an assertion that such a process was occurring somewhere else would be an error um so uh another feature so we have multiple inheritance uh so brain cancer in addition to being a kind of cancer is also a kind of brain disorder um both of those uh are uh are diseases so um finally we have in addition to physical and temporal things we have what we call information based things uh and uh among those are findings or these assertions that we might curate from the world um we have kinds of finding we have a we have a restricted set of of kinds of findings we add to it as needed um but it all falls within a unified model so one of those kinds of findings might be a Disease Association finding um and uh here's an example of something like that so we might see uh we might have a finding that although it isn't structured as uh as simple text like this we can generate uh text reflecting the meaning of this finding um some mutant human gene um in this case msh2 is the name of the gene um and we know the particular mutation and there's this is a very um this is a particular jargon that you'll be familiar with if you're familiar with genetic genetic mutations but suffice to say that this describes a particular um genetic variant some particular mutated Gene um has been observed with brain cancer in human um there might might be additional metadata telling us more about um who said this um how many humans um there but this by itself is is already um a kind of assertion we would model this um as uh an instance itself um of Disease Association finding so um the uh type of this uh instance is is Disease Association finding um but within that instance that instance is built of instances of um a variety of other classes right so um there would be an instance of brain cancer that was used uh for this finding it would be used uniquely in this in this finding um so statements themselves are objects in the ontology both statement types and statements themselves in order to uh put put this together as I uh mentioned earlier there are a bunch of uh different sources that we or certainly a bunch of different kinds of information that we need to be able to gather uh and we pull that from a variety of sources and we have um a variety of techniques that is not as as uh a long a list as the variety of sources but we uh we do encounter different uh kinds of workflows that we need to implement uh depending on characteristics of the source so um roughly one way to think of these is that uh sources range from being highly structured where we need uh we can pretty much automate the process because a they're structured and B We Trust um their structure um to things that are totally unstructured uh and I'll give some examples of each and talk about how we deal with those so uh some examples of what we consider to be adequately uh and accurately structured sources are um these things called Entre gene or omim so Entre Gene is uh the National Library of medicine's database of genes um we focus uh largely on human genes um although not entirely um bantree itself is a resource that covers many organisms and uh we collect from the not only the list of genes and the terms that are used uh to describe those genes but a variety of other characteristics uh and that is um pretty much doable with um an in an automated process uh or we can more or less uh trigger a script on a periodic basis and stay up to date on uh on that kind of data um a slightly less structured resource but one for which um the structure that it does have is adequate to our current needs um is omim uh which is stands for the online mandelian uh inheritance in in man I think um database and this is uh this is a database at John Hopkins University that relates genes to disease uh there's a lot of unstructured data in in this database but um as I said the the elements that are most um useful uh or I should say it's useful enough um at the level which it's structured for us to be able to also automate import from something like that um so at the other end of the spectrum we have what we would consider to be completely unstructured sources so uh peer- reviewed articles the literature um obviously a very rich source um but highly unstructured um and um or if it has structure uh a given instance has structure the structure is inconsistent um which is just as bad so um for that uh we have a completely manual process and so this is uh obviously um expensive but uh it Bears um it allows us to curate very sophisticated content um a we have so we have trained um scientists so first of all we first of all we we point scientists at the problem but second of all we have a protocol that they um that we develop that they are trained on and tested on U before we sort of let them loose on this process um they have a curation tool that they use uh to align their work into CER of facilitate uh what they're doing with the ontology and so they when they are uh curating findings statements that we have told them via the protocol are the kind that are of interest to us uh they can map them or they can see what uh terms map to the ontology um and if something doesn't map but they know that it's the right kind of statement it also gives them uh a way to request a new term in the ontology so this is a sort of a um semi manual or semi automated relationship um between or or or workflow I should say that allows uh the ontology to to grow and adapt as needed um so that's it um so it that's relatively slow and expensive but we get um very high quality content and it's um really currently the only way to get uh that level of sophistication from the unstructured from unstructured data um and in the middle we have uh or somewhere in the middle we have a variety of sources that are called semi- structured uh and so for these we have semi-automated Imports so we have an automated portion of the import but it's assisted by some manual process um either we clean up the data on the way in we notice on an automated basis that there are issues or potential issues uh and um and sort of patch provide a patch hatch layer for the data before we import it um or uh we might use a process to understand at a basic level there here are the here are the documents from that database that we know are of Interest we know that the data in them um is not structured uh so that at least allows us to go ahead and point curators at those particular uh documents rather than the entire Corpus uh so um examples of those data bases one of them uh interestingly to me is the clinical trials database so this is uh a quite large database of over now probably over 120 maybe 150,000 records um and growing um at something like 50 or 60 records a day um every day um of clinical trials that are either currently open at some stage or have been um completed at in some way for some reason Reon either they were they were terminated because they weren't going well or um or they finished and they did go well or somewhere in between uh so this is um a database of records that are um retrievable in XML form and certainly there's a lot of structure there what we find however is that uh you know structure person um an allegation of structure I should say per se isn't really always enough uh and so this is a case where we see when we look at the data uh we see that the actual um uh text or the actual content uh in each of those elements um is not always that well controlled um certainly um there's a control vocabulary for parts of it for not for other parts um there may be errors uh there may outright errors or there may be um uh sort of more subtle errors we can detect a lot of those things by comparing uh effectively uh what we're reading from that data uh to what we already have in the ontology or in in the knowledge base um and so for something like that we uh we have this automated import process but then we maintain a long list of patches against the data uh to uh resolve issues that are there but that we don't anticipate are going to be fixed by the data owners um and um and who actually owns that data is another um sort of complicated story um another area of sort of semi-structured data are FDA drug labels so um FDA um drug labels um are generated when FDA approves a drug for a particular indication or indications and um and they do have structure uh but it's sort of at a very high level uh and in order to uh get the really useful stuff we essentially have to identify the the labels of interest and um and send them for um some kind of manual curation so finally we have um another uh kind of issue which are um what I'll call sophisticated but unreliable sources so um disease databases uh are an example so um we use disease databases not so much uh for individual observations uh but as sources to help us with our terminology and with the ontology structure itself um which diseases are type of which other disease and how are they related uh and uh in those cases um diseases being a sort of at a conceptual level something that are that's not really um very well defined uh it sounds obviously has a definition but there it's um somewhat flexible and it's certainly flexible and um changes over time in particular cases so uh not surprisingly because of that we find that um basically all of the databases out there including the ones that are uh structured as ontologies um suffer from a variety of problems uh that are difficult to detect at a minimum difficult to detect automatically uh and so those might include missing relationships duplication um outright uh sort of Errors of commission um inconsistency internal inconsistencies so uh in order to to deal with that we basically don't use those sorts of uh databases as uh sources for the ontology so the ontology structure and the class structure um is essentially uh manually built and maintained um as a result and we use uh things like these as guides um and they certainly accelerate uh and facilitate the process but we found that we aren't able able in this uh regime to just take even for a particular branch of the ontology or part of the ontology an external ontology and use it um as is uh we have found that it's U much more effective uh not to do that and instead to to use them as I say as guides but an essentially manually manual process um and so yeah we and I so as I say here we use um both um inhouse experts scientists um as well as internal consistency tests in the ontology itself to detect those kinds of Errors so we don't rely completely um on people looking everything over um but because uh we do we are maintaining an ontology we can see internal uh inconsistencies uh and those can be flagged for for Corrections um in order to um I also mentioned that the content has to be testable so uh we have uh a long list of tests against the content and uh those are also uh encoded within the ontology itself uh and so we have a list of different test types that are uh sort of pre-cooked as it as it were as well as uh the ability to write more sophisticated uh one-off tests so in the in that first category uh we have um for example we can restrict slot values um to a particular type so we we might know that the name of something has to be a string um or that implicit location uh is uh is meant to point to a class uh to that is to a concept uh rather than a string or or a number uh or um slot values that the find at the finding level should be pointing to instances um we restrict cardinality of slot value slots have to have we can say maximum or minimum or exact numbers uh we can restrict multiple inheritance so that in that way we can uh indicate that things like properties and temporal things or processes um are mutually exclusive there should never there should never be a class that can be both um so that sort of Sanity testing um and then uh we can control string values via regular expression matching um at the more slow more complex level we can do things like prohibit duplicate finding so if there's a finding with the same structure uh we can recurse down into those um into each finding and ensure that we don't have uh two findings that are that duplicate each other um we can test the consistency as I've uh U mentioned once or twice uh of class modeling using properties on the classes so um if we know that we have a class brain cancer and it's defined as a cancer of the brain we know that that should be a sub we can tell that that should be a subass of cancer and if it's not we have an error somewhere and it should be a subass of brain disease um and that's possible because um the definitions um of some of these classes with respect to each other or or the respect to other classes um there are over 9,000 uh constraints uh overall of U and they vary in scope we run those nightly over over the ontology and that helps us keep things uh clean so I will end there um I want to acknowledge a few people so um the uh Ingenuity was uh was co-founded by Ramon felano uh Ramon and uh Tony uh loer were the authors of the Ingenuity knowledge representation system so this is all um um implemented in icis we call it uh icis itself is implemented in Java um and um that happened um now 15 years ago um and so it's been developed since then Cory cook um does the develop um heads up I should say the development uh and maintenance uh of IIs um Sarah tanam is the is our director of content um and um K parek is the ontology group manager I myself and am in the ontology Group which is uh in the content part of of the company um and that's my contact info and I will take questions yeah yeah so you guys are not choosing or stuff like that right so so this is a frame based ontology those are um description logic um that I mean it's a different way of representing an antology um so for and so right so the short answer is no we're not uh and um but we found that uh frame based ontology um is much more useful for us um it allows us to um be explicit about uh what belongs to what rather than defining rules uh for where instances should go um and um yeah other questions yeah you use all um very carefully uh so uh so so the knowledge base um is basically used to drive uh a variety of software products for the most part uh and those software products have in common sort of as a as a common theme this this idea that uh we are allow we're we're giving customers tools to align their own data with a state of biological knowledge and in that way um get an understanding of what their data mean uh and um well it's like it's adding it's allowing people to get insight uh right so so it's a it's a tool to uh give people context sort of see um what their you know why a particular list of genes might be showing up in their assay as being important in some disease like we can tell you oh that list of genes uh is uh representative of a particular biological pathway so what you're really seeing is that this disease must interrupt that pathway in some way um so that's a that's a you know sort of a simple and older kind of example but we can also take uh a lot of this data put it together and um we use it for molecular uh Diagnostics so uh there in a variety of of different ways but uh the idea is to understand if you have a genetic test for someone who's uh got a particular disease phenotype uh and the and the genetic test comes back and the sort of the the raw data um at some level of raw data that others um can address more and have addressed more um but the raw data is um and be a list of genetic variants well which of those variants is uh is relevant to their disease which one should I be looking at and what might they be telling me about uh potential treatment options um for that patient so we that's an example of bringing to bear like sort of the state of biological knowledge uh to understand like what otherwise would be pretty dry data it's mostly research uh well no I mean it it includes research um and it's certainly used uh for example in the area of drug development U but um in as much as uh we talk about the the diagnosis case um or cases um there are a handful of use cases where it's directly relevant to patients so it's not directly available to patients it's it's um this sort of software for example uh would facilitate um a company's work a molecular diagnostic companies work so um if someone the workflow would go something like this um someone gets sick they go to a doctor the doctor uh orders um a genetic sequencing test the the test uh order is filled by uh by one of these companies uh and the and the company is charged with providing a report um but what the company gets initially more or less initially is a list of mutations and the doctor isn't interested in a list of mutations the doctor needs um a list of um mutations along with their potential significance from this mutation is irrelevant to this mutation is likely to be relevant to the cause and um is informative in terms of uh telling us what treatment options might be useful for that patient so yeah actually brings the question so um how the users actually ask questions to the system mention a query language yeah so that's how they they would write the query themselves so the query language right so from a from the perspective of users so we currently we don't make the the knowledge base available sort of in its raw form um externally um instead it's used to to power applications and so um in practice what we do is the uh the knowledge is represented um and maintained and developed um and icas uh but then we export views of it um basically in what amounts to tabular form we so we export the results of certain queries um those get fed into uh relational databases and those um because they perform more quickly um are used to power the the applications and so um however if uh you know depending on the on the use case if a customer were to come to us with a particular question that wasn't the sort of thing that an application would have the answer to um we can certainly write queries against the uh against the knowledge base and and yeah and do that in the course of development okay thanks very much