Devreal

SF Text: Jeff Lerman, Advantages of Modeling Knowledge in an Ontology @Groupon

SF Text: Jeff Lerman, Advantages of Modeling Knowledge in an Ontology @Groupon

Recording: SF Text: Jeff Lerman, Advantages of Modeling Knowledge in an Ontology @Groupon

i'm going to be talking about the kyogen biomedical ontology uh thanks for that introduction um so we've we've had a an introduction already to ontologies in general i'll i'll probably be repeating uh some of that but our ontology is organized a little bit differently so this will be an interesting contrast i think um so kayajin is a life sciences company it they're involved in a lot of different things they actually acquire the company that i have been a part of for some time which is called ingenuity systems and ingenuity uh was and the the same sort of uh a lot of the same people in the same location and all that are still developing bioinformatics software a handful of different software applications i'll talk a bit more about them specifically but what they have in common is the idea of trying to take customer data and allow customers or i should say enable customers to gain insight from their data and ingenuity was born in a time when genomics data and molecular biology data was becoming more and more high volume experimental techniques were just beginning to be developed where people were able to get tens of thousands of data points um at a time it could do whole you could do surveys of which genes were expressed in tissue and there might be 17 or 18 000 genes and you'd have the answer all at once and there was a an analysis uh bottleneck for the data and so there's a question of how do we uh show people the patterns within their data and um the kyogen on what is now the kyogen biomedical ontology power is the software that allows us to answer some of those questions so um and i've been uh my formal title is ontology engineer i've been doing this now at ingenuity and al qaeda for uh almost seven years so i'm and we have a group of uh at any given time like seven to ten uh people working on the ontology per se um so i'm going to talk about uh this refill i'm going to talk about what the ontology can do give a couple of examples uh where is it used in practice uh some of the requirements uh that the that we have to impose on the ontology for it to be useful to us um what it looks like how we build it and maintain it and how we maintain as part of that or as a follow-on to that uh how we have what we do for quality control so all these things are are critical for us so a couple of examples of what it can do um this is a simple example so i want to know about all the lung diseases that aren't cancers so that's a that's a query that i can pose to the ontology or against the ontology we get a little over 200 results and they range from things that you might have thought of naturally to things that um might not have been quite so easy to come up with uh depending on your level of expertise um so we can start with like bronchitis right so and and we have a definition for each each thing that's uh free text and internal but uh i'm showing snippets of those for convenience uh so bronchitis should be one uh infectious lung disease is it a different level of specificity but another concept that we would cover um somewhere under that we'd have fungal pneumonia um you might have like an example of a kind of poisoning that affects the lungs so um berylliosis um and a bunch of other uh examples here uh just to give you an idea of the sort of uh the spectrum uh so we say lung diseases that aren't cancers that suggest that that already we can categorize things in more than one way and do some sort of basic manipulation of those categories here's a more complicated example i want to know about all the active clinical trials for drugs that treat cancer in patients that have mutations in a particular gene so okay i get 175 results from that query and here's a few examples with the drugs that go with them they're across just five different drugs uh so what kind of stuff has to go into this we had there's a lot of knowledge that has to be brought together unified and in order for us to synthesize this answer so we need to know about what things are considered to be drugs we need to know which ones are considered to treat cancer so we need to know about drug approvals which is and specifically for drug approvals what the indications are for the approved drug it's always part of an approval we need to know about drug indications in some more detail because we need to know drugs that are indicated for people with with mutations in a particular gene uh so we need specific indications and we also need a way of framing those indications in a way that separates the genes so it's queriable for us and then finally oh yeah we need to know which trials are out there and which ones are active so and these data come from disparate sources so um those are some examples uh of of what we can do and i'll and i'll try and get into uh what we bring together in order to make that possible and how we do that so ingenuity or i'll say ingenuity more often than i probably should but really we're not chiogen so kyogen makes software to contextualize molecular biology and genomics data sets so uh there's different kinds of contextualization uh the one of the things uh that we do and that we've done for the longest time is to make software to align customer data to the world of relevant biological knowledge so you tell me there's a hundred genes that i know are involved somehow in this phenomenon that i'm studying that's what my data tell me i don't know what those genes have in common i don't know which ones might be the drivers for the process but if i layer that list onto a knowledge base where i know about a lot of pre-existing or where i can bring in pre-existing knowledge about different relationships that those genes have to each other i can draw those genes on a canvas and i can start to see oh here's a system that might be starting to reveal itself to me and that might help design follow-on experiments um so we also facilitate insight into existing uh biomedical knowledge without even having to add in uh you know specific questions about my own data or a customer's data we can through the ontology power inference uh across multiple individual facts that may come or we say findings since we're skeptical about what's a fact but we can power inference of course across multiple assertions uh from disparate sources or from the same source and so i've given i'm showing here the two software products that are sort of most most prominent uh in our portfolio there's ingenuity pathway analysis which is um does pathway and network analysis and that's been around for over 10 years now and variant analysis is a little bit newer and the idea there is to be able to come in with a clinical data sets that contain that are from particular patients with particular issues um we might come in with uh with their and basically analyze sequence data analyze their genomic data to try and understand uh what the most likely suspect mutations are that might be causing their problem or might be closely related to their problem so um and that sort of thing is turns out to be really challenging because you take a genome any random genome and compare it to sort of the so-called canonical genome and there's on the order of a million variations right which one is which one should i be focusing on right so by taking all a lot of pre-existing uh knowledge and and uh a certain amount of insight uh and overlaying that we can quickly filter out the things that are uh likely inconsequential or irrelevant and find that we're able to arrive at a much shorter list of uh likely suspect variations so um all right so we need biological knowledge to be organized in a way that's accurate um and we need to be able to by the way to test that accuracy um it needs to be normalized in at least a couple of different ways one is we want to be able to think of concepts not terms and so we want to be able to map synonyms to a concept meningitis has these couple of individual synonyms uh and we that's one kind of normalization and we also want to be able to map different syntactic structures to the same meaning so meningitis is a you know a single term but it can be defined in terms of other concepts it's it really means inflammation of meninges which some other which arthur some other author might refer to as the minix although hopefully not so this idea of having different ways of saying the same thing not just using different terminology but actually with with different structures is something that we have to be able to get past it has to be sophisticated enough to be useful so we have to be able to capture when we're capturing assertions from the world we need to be able to capture different kinds of assertions and we need to be able to link them together so that so that we can really get the inference power that we need so here's a i have a very simple and mathematical looking example we might have one finding that says tells us that a some proteins say causes some other pair of proteins to bind each other b and c and we have another statement that says b c binding increases process d well we want to be able to infer that what in this case might seem obvious that well that should tell us that assuming that the conditions are met for both of those statements at the same time that a should be increasing process d right but those two component statements are of different forms and so we need to be able to have a unified system that allows us to infer across them finally we want a system that is efficiently extensible right so and this gets us a little bit away from a relational database approach where adding new kinds of relationship types can be a real challenge we want to be able to do that quickly and easily and manipulate those those sorts of things so i have a short list of some relationship types that me we might want and and have uh so physical part of something is a physically a part of something else implicit location is something that we use and we'll see again to talk about uh the location that a process might have um that as a as a reader or as or as a domain expert you would say oh of course uh but the ontology has to have that modeled in so that we know that that a particular location is implied by a particular uh process that's happening another example has mutations so some gene it might be asserted to have a particular variation or mutation um so yeah do you manually determine the extended relationships do we manually determine yeah those extended relationship types like this do you manually determine them can you automatically determine so um we want any examples of the relationships or when you want to figure out what extensions you need to make can you automatically discover them or you have to sit down so far that's a manual process so right so the question was uh can we automatically determine if we have we have some way of automatically calculating uh when a new relationship type is needed um and the answer is no we might be able to come up with such a thing but we don't have anything in place to do that now you speaking in general that there's so yeah i mean i can't speak in general to whether anyone has come up with such a thing but i mean i can say for at least for for this space that um yeah we we still have to do that manually now we can we can infer the need for an individual assertion on an existing relationship type and we can sometimes infer but still manually the idea that a new relationship type is needed for example when we start to see violations of some rules that we can't resolve by sort of using our existing relationship types we realize oh actually we have some for example we might have some relationship type that really is encompassing multiple relationship types and ought to be ought to be split uh so we encounter that sort of situation um that's maybe as close as i can think of that we come to an automatic process for deciding when a new type is needed um so we use a frame based ontology and i'm going to talk a little bit about what that means so first of all uh we have concepts that is to say kinds of things and the kinds of things can be any kinds of things now um we um we model the things the idea the concepts of the things directly uh so we do we we we are not really at a removal we're talking so the previous and jay was talking about topics um we talk about the things themselves uh so we would have a concept for organism a sub class of that or a sub type of that would be mouse or e coli there's other physical things like cells but then there's other things as well and we'll distinguish those things and i'll show more of that in a moment but so diseases are obviously not physical things um cancer is a kind of disease and then in addition to processes and physical stuff we also have concepts for properties and we distinguish that as well so advanced benign narrow motile those are all you know adjectives basically yeah well sure you can think of it that way we would prefer to think of it as more general concepts so and instead of talking about chinese restaurants we would talk about the concept chinese restaurant um there could be and then there could be more specific uh a more specific concept to uh cover um chinese restaurant and you might even have yeah and on and on uh so um we don't have a bright distinction between very broad topics and very um and much more narrow topics um concepts but we find that it solves a lot of problems if we consider the the concepts or the classes to be representing the things themselves um and that might become more clear later on but i'm happy to answer more questions about that as well um so okay so physical things uh temporal things or processes uh properties and properties of course can be um subclassified in their own ways to what they're what they are applicable what they're applicable to as well as uh sort of assertions themselves and kinds of assertions also have a place in the ontology so in addition to concepts we have relationship types as as noted so is is the most prominent kind of relationship type um but then we have um many others i was i was looking earlier and we have um about a thousand different relationship types uh in use in the ontology right now so uh physical part of is is uh one example of so um i guess i think from a topic point of view for example illinois might be a sub-topic or sorry chicago might be a sub-topic of illinois um from our perspective uh chicago has a particular kind of relationship to illinois which is physical part of and is obviously distinct from an is a relationship it's a physical part of and then this is a more recent one definitional agent so bacteria are the are the implicit agent or definitional agent for bacterial infection right so so we would consider it back we would we might think of bacterial infection as being a kind of a compound uh concept and it can be defined in terms of its parts right so bacterial infection is infection by bacteria it's a kind of infection but it's not a kind of bacteria right there's a different relationship there so um so concepts are encoded in classes what we call classes and relationship types are encoded in slots which is a and both of these are terms from sort of the standard way of thinking of frame based ontologies in addition to those two kinds of things uh we have instances and i'll well i should talk about that now so um instances um for us are not related to how specific something is but rather a particular reference to a thing gets its own instance right so a statement about mouse that we that we encode would not refer to the class mouse it would refer to an instance of mouse right uh and that allows us to to assert all kinds of details that we obviously are not appropriate for the class itself so um so we have classes slots we have instances and then we can also encode scalars so strings numbers booleans there's a few other kinds that i won't talk about here and the first three things are what we would call frames those are frames in the ontology um so uh and scalers are not so a frame here is um it's this is not a use of the term that i think we're used to hearing it doesn't mean like a box around something but it's rather think of it as a super class if you will for the idea of classes and instances and slots anything that's a frame can have relation okay we can say things about a frame when we can assert relationships among frames so we can assert relationships between instances between classes or even we can have relationships from one slot to another right we can talk about the relationship between two kinds of relationships um so uh that's the that's the very sort of fundamental structure that we work within uh for frame based ontology and for ontology um so here's a little bit about what this looks like um we'll start with um two concepts and anatomical part and physical thing so anatomical part has an is a relationship to physical thing any anatomical part is a physical thing below that we might have the brain we might also have the hippocampus now you might notice that there's probably a relationship between those two but it's not an is a relationship right they're both anatomical parts we need a new kind of relationship to indicate that something is a physical part of another so i'm representing that in what actually does look like blue good um and um so that's all so physical things are under one we are under one sort of uh tree in the ontology um and i should say now because i don't think i i'm gonna say it otherwise that we cheat a little bit so the for us superclasses or the superclass relationship is synonymous with isa and so basically we get one slot type for free or one relationship type for free and so we assert superclassing what we mean to say something is uh something else um all the other slots have specific names and are specifically asserted as in a slightly different syntax so at the top of our of the resulting tree we have a very small number of roots we have physical thing we have temporal thing and we have a few other things that i'll that i'll talk about so temporal thing something that occurs in time we include under that among other things diseases we consider diseases to be processes this disease is sort of a made-up concept to begin with so this is this doesn't always work as well as you might like but it's the best that we've been able to do it works most of the time and then we can you have more specific cancer brain cancer so brain cancer can have this implicit location brain so this is the same thing as i was a little similar to what i was talking about earlier say i'm talking about defining compound concepts in terms of their component parts right so and we find this to be very useful in terms of maintaining quality of of the disease branch and of the temporal branch in general is i'm keeping track of which things are implicitly in a particular location that's one of the pieces to our puzzle we support multiple inheritance so brain cancer in addition to being a cancer is also a brain disorder both of those things are diseases we prohibit the only thing we prohibit in terms of super-class relationships or is our relationships are cycles so we we don't have we don't have any circular uh relationship relationships um on is us slots although we don't formally prohibit it on another slot so that's sometimes we we need i in a few cases we might have rules like that but the ontology uh engine itself doesn't uh prohibit it except for is a slots uh other than that anything goes in terms of is a unless we um define a rule um to indicate that we don't want to see certain kinds of things so for example um we probably don't want it we definitely don't want to see any properties that are also considered processes or any physical things that are considered temporal things those are those are mutually exclusive and we have ways of enforcing that so i mentioned that assertions themselves get classes in the ontology so we have something called information based thing findings go under those and then we have different kinds of statements that we can make and each kind of statement gets gets a class so one of those classes is a disease association findings already fairly specific an example might be so a particular disease being associated with a particular mutation so we see here a particular mutant gene and this cryptic thing here is the description for a particular mutation has been observed with brain cancer in human um so to make this this statement and i'm not going to break it all out um here but a statement like this would be modeled in terms of instances of each of the component parts so for example there will be an instance of brain cancer here there will be probably one instance of human for the location there will be an instance of the msh2 gene there'd be an instance of mutation with some details including this description etc and all of that will be hanging off of sort of a backbone that was an instance of disease association finding so and all of that lives in the ontology all right the short answer is no we have a lot more information so i mean there's different kinds of information that we have about diseases one kind of information that people are very interested in is what mutations will make me sick and what mutations will what mutations correlate with treatment by a particular drug a successful treatment i should say uh with a particular a particular drug or class of drugs uh so there's a lot of interest in observations that um show some uh co-occurrence and correlation is a stronger term and causation is the stronger term yet but we start with co-occurrence and in individual cases and then we can build up to correlation and maybe we can infer causation if we know enough about the mechanism um or we can learn about the mechanism uh so but there is a lot of interest in these particular kinds of observations yeah right so the data come from several sources and i'll talk about that in a few slides so um let's sort of put that off momentarily but you're totally right um i put that there just for you no that's correct so uh type of really should point to the thing that it's a type of you're talking about this arrow here so the the the question was why is this arrow pointing the wrong way and the answer is because i made a mistake so this one is pointing the correct way right so typeof means an instant is a relationship that is reserved for relating instances to classes and it indicates what it basically is the only is the primary thing i should say that defines what an instance is right um anonymous instances can live in the ontology but we don't like to see them um so um and they're usually mistakes um so yeah um okay so i think i've said this but so statements themselves are objects in the ontology both kinds of statements and the statements and the statements the individual statements so data sources um we have there's a wide variety of data sources that we use and i'm definitely not going to talk about all of them but we have a variety of them and we have and we need to keep up to date with them for them to be useful to us and we have a variety of techniques to get data from them into the ontology so um and i've sort of divided them here into their own little ontology so uh we have structured sources that we trust pretty well and for those sources we can write scripted imports and so one of the things that we do is maintain scripts or you know some relatively small programs that parse these different sources so like entree jean is the name for the national library of medicine that's the the nih's database for genes omim is the online mendelian inheritance and man database this is hosted now at johns hopkins university and it's a database that is now online but goes back to the 60s and is basically descriptions of relationships between genes and diseases and and particular variants and diseases um so this first one is very highly structured omim is a little bit less structured but still has the the parts of it that are useful to us so the are the structured parts and in both cases we write scripts that go through all the steps of parsing first of all we have to decide what kinds of statements we want to bring in from these sources and then we have to map each of their terms to concepts in our ontology so we in that sense we're using the ontology or the parts of the ontology as a dictionary um right and um so anyway we have scripted imports we run these periodically we make them incremental so that for efficiency we're only updating the parts that have changed uh we keep track of where data came from when we got it what date was asserted on it in the original source if they gave us a version or a date etc on the other end of the spectrum and there is a spectrum is the literature so we also curate a lot of stuff from the literature and of course the literature by its current nature is pretty unstructured right we're basically dealing with pros uh interspersed with figures and tables and for that we need a manual process and so we have a whole process to train people to go into to go through articles and extract the statements that are relevant to our needs so we don't use the tagging tools from nlm but we do we prioritize the articles that we want to go after partially on the basis of the tags that are applied by nlm so articles for practical purposes any article that we carried has to come from has to have an entry in pubmed the public repository of biomedical literature that's hosted by the national library of medicine medline is sort of a refined version of pubmed but we don't wait around for thing records to graduate to medline we're in a hurry so sometimes we're lucky enough to have uh the the more more tagging that's been associated sometimes we we don't um so we want to reduce our dependencies on sort of outside uh sources especially when it comes to the literature so that we can curate something with as few bottlenecks as possible but but we do utilize it when it's there to help us along your first pass right um so tagging we find and talking is useful to us for two reasons one we want to identify articles of interest in the first place um that we might ever want to get and two we want to identify the priority of an article which ones we want to get first just assist first pass so i can't get too deeply into some of this but i will say um right i mean it depends we have a couple of different processes depending on we have we have several workflows actually for manual curation um what some tools but not necessarily those tools we do want to help the curators as much curators as much as we can but at the end of the day they're they have to be making they have to be expert enough to make the right calls in terms of the individual findings that they're extracting in the model um that they're using uh so um yeah that i guess that's why i'm framing the sort of the tagging question as i as i am we wouldn't use tagging uh in place of curation right right yeah to save people time and to help them understand uh you know it's like oh we're giving this paper to you we're we want you to carry it out all of the relationships between mutations and diseases maybe you'd like to know which diseases we think are in this paper and which genes right so we'll we'll take a good stab at that um right i mean parts of the paper guess it gets a little interesting right right i mean i think the closest we come to doing that sort of thing is to help us understand more specific topics right so if we've identified say a certain combination of topics that are probably discussed in a paper that can help us understand that it's in some more specific topic that that is that that represents that synthesis um but in terms of that's that's probably as far as we currently go um so so there is this manual there are like i said several manual creation workflows these are tool assisted and the tools are informed by the ontology uh so in addition to sort of whatever pre-tagging we're doing for literature articles when an article is actually in the process of curation the curators are using tools that already know about all of the concepts modeled in the ontology so if they enter a term that they see in the paper we can tell them oh yes we know about that term that's a valid term at the same time they have the opportunity to make a term request so if they're not convinced that the concept that they or the or the term that they um need in order to capture that finding is available in the ontology there's a built-in workflow for them to make that request and that goes to another team and that some other team does the modeling and the addition into the ontology all right so somewhere in the middle here we have what i'm calling semi-structured sources so that's things like as it turns out the clinical trials database so the clinical trials database is a structured source it's available the records are available as xml documents and they're pretty pretty nicely structured but we find that first of all they don't go to the level of detail that we always need and second of all sometimes they're a little bit cavalier about how they how things are struck not structured i guess but but what terms are used within the structured data um there are particular sort of systemic reasons why that's the case for that database they're not that important the point is that we can uh we can pull data out we can import data into our ontology from clinical trials with an automated process but it needs a little bit of an assist and so we have that model as well where we have people doing a certain amount of pre-fixing we main we might maintain a uh a list of fixes that we want to persist um into that data source we you know we found an error we want to fix it but next time we do the import we don't want that error to come back so we we maintain like a patch list essentially now is probably a good time to mention that in our process when we're doing these imports from structured databases it's not unusual for us to find problems um and because we're synthesizing data from so many different sources and because we have such a strong need to for quality we find that we're very well positioned we're running our tests to find issues that this the authors didn't know about and uh so we actually developed relationships with the owners of these data sources and basically they fall into one of two categories people who are interested in the feedback and and very happy to have it and fix things right away um and people who are you know not disinterested but less responsive um and so you know we adjust our behavior accordingly um i it's easiest for us if someone can just fix the stuff and and then we'll pick it up next time but we also maintain workflows that allow us to keep track of the changes that we would need to apply if we don't expect them to be committed by the by the owners so um that's that's so tempting but no we really can't do that i'm sorry what sort of errors we've sure um so one that i can talk about because they're very responsive is clinvar so it's not mentioned on this slide but clinvar is the is a fairly new database being developed by nlm and maintained by the folks over there um turret and it stands for something like clinical variations and so it's basically they're capturing records of these observations similar to what i've talked about a couple of times of genetic variations and their correlation or their co-occurrence with disease particular diseases um sort of both so so basically it's individual individual cases but they are sort of a secondary source and they they themselves import data from other sources including a lot of clinical labs but also including some other databases so they themselves are consolidators and we in turn um import data from them uh but we find for example that uh so so the description that i showed earlier for a mutation this is in a particular language there's a syntax here and there are syntax rules that you can apply um and writing those rules in a in a way that you can run them automatically it can be a little tricky but we've gone ahead and done that because it's so important to us and so we can at a glance say to the clinvar folks hey we just imported you know 200 000 of your records and these three have invalid syntax could you have a look and they're like great thanks for letting us know and so so that's that's one kind of example but then we'll also have um more difficult cases like in clinical trial records um there's a place for the condition and that can mean the can the disease that's being treated or they were that they're that they're testing the drug against right so clinical trials are basically about treatments and and uh and diseases for the most part not all of them um so the condition might mean oh this is the thing this is the disease that the drug ostensibly treats and and that's the hypothesis that we're testing or it might refer to um one of the entry criteria for the trial right pregnancy smoking um right so we're not trying to cure pregnancy right but um that's that's an example of um maybe sort of a database issue where most of the time it really does mean the indication for the drug um but not always and so we have to deal with that there are a number of other examples is that data right although we see some of the latter we don't see i don't know if we see experimental errors um one of the things that we find in the literature though is so for example we might have an author describing a dna change and the corresponding protein change we'll capture that as a single as part of a single finding and we have all kinds of tests that we run against our data including can that protein change come about from that dna change sometimes the answer is no and so in a case like that we have to go back and figure out what the source of the error is sometimes it's the curator right our reader typed something wrong sometimes it's the author of the paper nobody ran their paper through um it's an analysis engine right there was peer-reviewed it was reviewed by other human beings uh so um these kinds of errors do slip through we're all humans and so we we do find uh mistakes like that that are you know conceivably more insidious i'm sorry it seems like you could convince the scientist to have somebody do the data entry for you um things like that would be really cool we're not we're not there yet and it wouldn't be us um but the idea of having like a public repository where certain kinds of um findings in uh in published literature were uh were recorded and that repository could be responsible for doing those kinds of tests would be great um the one example that well the example that i'm most familiar with in the in the biological sciences that does do something like that is the protein data bank so if you take a if you're a structural biologist and you solve the structure of a protein especially by crystallography you end up with data files in in one of a variety of formats and one of the sort of basic requirements of the field is that if you're going to publish that structure you have to deposit the data in this repository now when you deposit the data there's a series of tests that are run against the data to ensure that it makes sense there's basically data integrity and if it fails those tests you can't deposit and so if there were something like that uh for in the area of genomics it would be great but it wouldn't solve the problem that there are decades of literature out there that people are still interested in that existed before even if that repository were created tonight like right we still have to be able to capture data from the past uh so we don't it would help us going forward and it would help a lot of people going forward but we still have to have a process like this all right uh oh okay so finally something i haven't mentioned or what we call what i might call sophisticated but unreliable sources so like diseases are a particularly difficult problem to solve because they're kind of synthetic concepts so um or made up maybe is is fair how you classify diseases and what disease is a another kind of another broader kind of disease is a matter of some debate and fluidity right opinions change as our understandings of disease mechanisms change so there are different databases out there some of them are ontologies that attempt to classify diseases but we use them as guides um and they're very helpful but we don't find them uh sufficiently good to be able to import them in any kind of automatic way and so things like the disease branch of our ontology are basically entirely manually curated so that's a thing right so like i said these are still useful as guides to locate and correct the errors we have resident experts so you know we might say domain experts and we have a variety of internal consistency tests which i keep talking about so within the ontology in addition to the physical things and the temporal things and the information based things the findings we also have a branch of the ontology just for tests which also exists as assertions inside the ontology or objects even inside the ontology so and we have a variety of types of tests or constraint types and we have the ability to make more on the fly so and these and some of them are built in and some of them are what we call roll your own and so for some of the simple kinds of tests regardless of where they came from we can restrict slot values to being to to be of a particular type so values of this relationship or the destination for this relationship might need to be a string or a class or an instance of some particular class we can do the same thing by the way for what we call the slot domain where a slot can be attached where it can come from um we can restrict the cardinality of slot values so the number of values that a given slot can have when it's attached to a given frame and we can restrict its exact number or its maximum or its minimum we can we look at that so for example every class has to have exactly one maybe it's easier to say title so that doesn't mean it can have only one term that maps to it but there's one term that's the name and everything else we call synonyms or some other kind of term so what else we can restrict multiple inheritance so remember i said properties and processes or properties and physical things are mutually exclusive um we need to be if if that's going to be true we need to enforce it with a constraint so these are all tests that have to be written into the ontology we there's a regex engine so we can control string values with rex matching either positive or negative tests and then there's some more complex kinds of tests we can prohibit duplicate findings but that involves digging into the structure of each finding and looking for ones that have this the same structure maybe not the same exact assertions but the same structure so what else we can test the consistency of class modeling so this becomes very powerful for us and this is one of the ways that we maintain the integrity of the of the class structure so for example brain cancer which is defined as cancer of brain that is to say in terms of cancer and brain should have a number of characteristics we know that it should be a subclass of cancer because it's cancer of something and we know that it should be a subclass of brain disease because oh by the way we have a definition for brain disease that it is disease of brain right so and we also know that this is disease as cancer brain cancer is a disease so we can we can infer across the ontology to look for hints basically that we're missing a relationship or sometimes that a relationship is incorrect so we have over 9 000 individual constraints in the ontology we run them every night and we we deal with the with the violations we classify them in terms of how critical they are we don't treat all of them as priority one but that gives us the ability to know where we stand and in addition to that we have a system set up such that any change made to the ontology uh triggers running of any constraints that are likely to be relevant to that change uh and whoever committed that change if their change seems to have caused a violation gets an email and so that happens within the hour so and that sort of quick feedback process allows us to reduce the level the error rate overall we're not in a position where we can instantaneously uh run tests at the time of of assertion uh but we're getting closer um all right so with that i'll finish up the ontology so i haven't talked at all about the engine that drives the ontology but the ontology is um run or it lives inside um a thing that's called the ingenuity knowledge representation system it's our own internal software um and that software has been around now for um really as long as ingenuity has been around it's developed and maintained currently by corey cook and i'll mention a couple other names sarah tannenbaum is our director of content and for a long time managed the ontology group and kaushal pareq is our current manager coordinating the group um and i can be reached here and i'll take any last questions one thing i was curious about the touchdown but didn't really go into much detail on is is taking different ways of saying the same thing and so my question is how do you what process do you have for things not manually period to say you know meningitis and medics are so um i would say it's the common we automate it as much as we can and then the rest is is manual so and there's different layers to that we have different mechanisms for asserting those definitions uh on on compound classes and then those definitions in turn are used to unify the meanings at another at another layer and that layer is pretty much fully automated so we have what we call the root ontology which is where we store all the meanings of stuff the the findings live in another part of the ontology um but then somewhere in between there uh there's another layer where we store all of these um we call functional annotations uh so the the meanings that uh that we might have and if more than if a meaning can be asserted in more than one way uh and we learn that from a definition or something like it um they get joined in that layer so is that is that helpful yeah okay yeah so we don't really know the medical things at all so i'm kind of kind of dark as to what what's in these papers where they like i'm assuming well that's that sort of content certainly is out there and i would say that uh i guess i should make the point that we focus on what i would call qualitative rather than quantitative data that we're capturing we're capturing relationships and sometimes there's a need to capture quantitative measurements but it's surprisingly infrequent you would find links you basically take in these links in a chain and then you assemble chains on requests yeah and for some value of request that's right now i mean that's not to say that we don't capture quantitative information at all for example someone's doing a population study um and they find that you know of the 100 people they looked at with this disease 70 had this mutation and the other ones did not we want to know those numbers um so um we'll we'll capture that kind of thing as well but we but the statistics are um i would say there's more statistics statistics happening after the ontology than sort of encoded in it right so we might do statistics on the population of findings and learn something that way um but it's usually too much detail and it sort of gets it to a level where it's it's harder to do useful inference if we're capturing the precise statistical measurements from each assertion in each paper so we generally don't do that all right you