Devreal

SF Text: Jeff Lerman, Q&A with Alexy Khrabrov @Groupon

SF Text: Jeff Lerman, Q&A with Alexy Khrabrov @Groupon

Recording: SF Text: Jeff Lerman, Q&A with Alexy Khrabrov @Groupon

thank you hello everybody I'm Alexi krabrov the organizer of SF text and here we are on location on Groupon we have a Meetup about ontologies and here with us we have Jeff Lerman staff Anthology engineer at Kaya gym uh who works on uh biological uh knowledge organization and I'm first curious can you tell us a little bit uh what is the knowledge domain which where you're building this ontologies okay so the kyogen ontology which formerly I should say the Ingenuity ontology um covers uh biology broadly is specifically for use in analyzing genomic data clinical and research data as well as more broadly molecular biology related data so and the original impetus for setting something like this up was probably you know 15 years ago as the just the volume of data started to get really large people started to be able to do experiments in parallel where they could acquire a large amount of data in a short amount of time very easily and the challenge then becomes analysis and so obviously we're still in that space today where we have a lot of data and Analysis is is the bottleneck the sources of the data are different and the volume and the data are different but that's that's why we do what we do so is this anthology of mainly for humans so is it basically organizational this knowledge in a human readable form uh he was supposed to computers right but because you can use indulges in different ways sure uh so the ontology is rarely used directly we don't make the ontology available we aren't currently um uh we don't currently have any business models we're making the ontology we're directly available to customers uh we do have internally ways to visualize what we have and to we have our own query language and our manipulation language but the ontology is really serves as the underpinning uh to other application software that we sell and so uh we need it to be computer readable certainly computer usable but at the same time it's serving uh it's not very far removed from different applications where it has to be human readable okay okay and so the the so the end user is still human it's not as you know like it's not like you use this Anthology as a way to automatically cluster things necessarily like you you have a human asking queries right and receiving the data back that's right sort of yes I mean it's interesting the ontology sort of um kind of sits at the core of the Hub of of what we do because on the one on the one side we're in the business of curating data from uh disparate sources from databases they may have a different organizational instruction and we've got um all the which are you know but they still have an organizational structure to fully um a sort of open-ended text from from the literature and so we do need the ontology to help the people who are importing that data um to to assist them in the process of classification and organization right so the ontology is both the system by which we organize data and it informs the tools that we use to allow us to have that organized data and acquire that data on the other side of course we want to expose that data in various forms and various circumstances through our application software to users and so of course you know we have human users at some point right but the ontology is non-accessed directly even by the applications uh we we take its organizational structure or output we export that in some format that is then in turn used by the application so there's some layers in between a few layers in between the end users and the ontology itself but it's still you know very strongly determined what our applications can do got it so there is a lot of data in in you know the people called bioinformatics or genomics and we have actually several startups uh even as a scholar which I also run and uh I was amazed by looking at you know genomics databases how much text is there how much artifacts so you know genomic sequences themselves various papers talking about them various calls right like this paper so and I used to work at present at the bank a long time ago and you know they had their own you know and I looked at Medline and so forth and I was really Amazed by the by the sheer volume of all this data and and and so it seems that everybody has their own databases their European database that OS databases there are all kinds of uh of things so have this uh that you're talking about are they unifying all these data sources for instance you know sequences themselves and and papers about sequences can he plays both things and you know on the same Anthology sure so uh you know we we certainly um Define our limits and and so that's one thing actually that we don't do in the ontology itself is to deal uh with sequence data um we do deal with variant data quite a lot and so that's becoming a more and more important as it's become cheaper and cheaper to sequence uh you know DNA and and even sequence whole genomes um which is sort of just ridiculous that we're at the point where we can do that yes uh but uh you know even a single genome results in an enormous amount of data even after it's Consolidated and the raw data you know much more so uh so uh what's of interest is to be able to look at the differences the the variance uh the genetic variants and to be able to learn something from uh the presence or absence of different variants so so we don't process whole sequence data we don't and we certainly don't store whole papers but we extract um I would say useful information what we consider to be useful information um statements um more than so qualitative more than quantitative data statements and relationships from all these sources so from from the literature certainly um and there are parts of the literature that we focus on very intently um and then uh there may be there are for example databases that focus on genetic variants and we'll import those and what and so we do have a task of unifying um a variety of different data sources so as you mentioned there are American and European databases for uh for genes and and other genetic information um I don't know about a competitor to the pdb to the protein Data Bank I think that's a fortunately centralized I mean based on my knowledge from 10 years ago right and it's been a while so I also used to be a protein structural biologist and um but I also haven't thought too much about the pdb for a while now but anyway there are obviously different databases covering some of the same information they complement each other to some extent and we do our best to unify information from those databases we want um you know we can't stand alone uh people coming to us with their data are their data are speaking a certain language right they're using certain identifiers to refer to Concepts we need to be able to map those identifiers to those same Concepts so we need to be able to take the refseq for example for one of the uh one of the American database identifiers those those IDs and we need to understand those but ditto for the European ones um and that's just for for genes and for genetic sequences there's a variety of other things like that interesting so uh let me give me if uh throw some idea on the table which I discussed recently with uh several folks in the space so we have you know several startups doing genomics and also amp Lab at Berkeley is using spark to to help basically with sequencing and and essentially there's several there are several communities both industrial and academic which unite to fight cancer through this kind of open genomic research and there is a Global Alliance of several forces and so you know being a software engineer myself I'm always thinking that you know open source uh approach is superior because you can you know put something on GitHub and kind of you have a good readme file and good example and test you know uh then then you will have a lot of developers who will Who will jump on it so so we have these conversations where we can you know we're thinking uh how can we package this knowledge genomics for computer scientists and what kind of resources we need to enable open source collaboration right so we need some kind of data which is referenceable and some common understood format and we also need to explain the problems uh in genomics to to developers but because we have a lot of smart developers in the space they prop can probably uh do a lot of advances so I'm wondering so you're kind of you're organizing all this knowledge right and um are there any insights you can share is the rename and there are any ways for the community to leverage this information and kind of quickly understand or Point people to the right pieces right so is this something which which you know software engineer can do as a hobby in their spare time or is this something you know on the professional researcher can do uh what is you thinking about this so the idea here is is you know how much of what we learn or what we what we organize can be um made available so that in a public domain in the public domain or if you know I understand you guys have a business model and you know we make these products available but how much of this can be uh can be you know you know why by open source Community for instance looking at the open available date this is something which which Community needs to do uh is this you know other ways for Community to organize this knowledge and make it available to computer scientists kind of to quickly educate themselves well I think yeah I think there's opportunities there um and frankly I don't think that that's something that we take advantage of or we really promote right now because as you pointed out and you know we have this in common with a lot of people I think we're a private company and the primary uh you know all primary energies are devoted towards making things that we can sell unfortunately um on the other hand uh we benefit from the existence um of public databases of genetic data and among others I mentioned the genetics that's a big Focus lately but there are certainly other public databases um that that we use um and um you know we do a couple of things so we we take advantage of those of those databases but then we also put a lot of energy into curating like I said unstructured knowledge and essentially uh exposing the structure right or exposing exposing uh putting putting a structure around the the meanings um so that ladder stuff is I would say probably going to continue to be pretty proprietary because uh that labor intensive kind of work that just requires a lot of money that's where uh the sort of private domain probably has an advantage over the public domain um meaning we know we need a large number of very highly trained experts doing this stuff and it's just expensive but uh for the public databases that we use we um they're more useful to us if they're higher quality and so we regularly provide feedback to these databases and they get to know us because we because we are so focused on on quality and testing and integration we notice issues that they frequently don't have the resources to notice and so we do provide feedback in that way and that sort of thing if there were a form for it might be even more valuable if we could say hey we have some suggestions as to how data could be and we're happy to share them with the world because everyone benefits right uh we have some suggestions as to what people should be paying attention to uh when they're developing databases when they're organizing data when they're when they're collating data um and uh and here are some techniques that people could use uh to um help that along and to make that feasible um so I think that sort of uh interaction uh I think there's room for that sort of thing makes sense so but I'm also curious you know so you said you're a biologist right so what brought you to from biology to basically knowledge organization which is essentially you know a computer scientist kind of domain it is it is I mean we're we're an interesting we're an interesting space uh you know in my group we we certainly are are doing data science as it were and we don't have a we don't use that name but people people do um we are a mix of people with a strong computer science background and people with a strong Sciences background I would say most people in the group are PhD level uh scientists either in chemistry or biology okay um and with uh varying uh levels of exposure to the clinical side but probably more on the research side um and uh it we're basically we're a group of people who are interested in um maybe stepping back from the uh what you might call in the business world the individual contributor level when it comes to uh scientific knowledge and more uh sort of aware and interested in the idea that hey there's all this knowledge out there but it's not that useful if it's not accessible and so much of it is inaccessible um either because it's not well um organized to begin with and that's how a structure to begin with OR because it's just not um integrated into a unified system or any kind of unified system and so there's knowledge that can be exposed without doing anything in the lab just by sort of putting the pieces together that are already out there if you can reveal those pieces and get them into one system so uh that that's intriguing to me I mean I came from a place where I was a graduate student and then a postdoc I guess like we said molecular biology specifically in protein structure and I did sort of biophysics biochemistry which is a lab I was in the lab I've also but for a long long time I've been interested in computers and sort of playing with them and so that's always been kind of a of mine and at some point the idea of using computers as a tool to help organize knowledge and make it more accessible and sort of do great things with it became more appealing maybe than the toiling at the lab bench and so I kind of turned that corner and so I haven't it's so important to me to stay connected to the sciences and and to the science that I was doing and I'm and I am um in the you know where we are we're still thinking very much about the science aspect of things um otherwise there's no usefulness to it right because it's not abstract knowledge like you need to understand the domain we need to have them really good domain knowledge I mean we need to understand what we're doing or else and that's what gives us you know some Advantage um is that on the one hand we understand um how to sort of reduce data and how to reduce knowledge and model it in a useful way and it'll allow us to do inference and calculation on the other hand we understand the knowledge itself and the bits and pieces and so uh that hopefully helps us not to make silly mistakes and also to uh sort of recognize the opportunities so I'm going to kind of come back to this you know open source uh Community to fight cancer because I was really you know struck by how much uh opportunity is there and how little is uh of this you know centralized organization there is a lot of organizations trying to unify this data but there is so much like it you know from a computer science scientist point of view is extremely fragmented it also seems to me that you know the result of great Sciences people but they're not necessarily Google level computer scientists so they because you know they devote their efforts into discovering the primary knowledge and there is not enough folks like you who you know piece it together yet right and so uh so I'm wondering is it even feasible is it possible for a computer scientists like me who does not have a formal you know molecular biology training um but is curious and kind of can understand almost everything if you read it many times uh presumably um Can can we can we present the problems of fighting cancer indigestible pieces can we decompose a problem and can we also you know take all this knowledge social mapping and present a computer scientists and digestible pieces and kind of map a little path to them so somebody on over over a course of time a community can kind of make different uh efforts and kind of together uh write a lot of good call to eventually you know defeat cancer through Community you know crowdsource programming is it is it a possible possible activity uh right so I mean I I think when we when we use terms like defeat cancer we kind of we're at a very high level and and um you know sometimes that can you know but but there are lots of uh foreign I mean you know so one of the things people are really interested in doing now is uh finding ways to leverage genomic data genetic data uh to determine what the best what the most likely uh good treatments are for individual patients yes right so we talk about personalized medicine that's a lot of what people are talking about yes um so in order to do that we basically want to be able to take advantage of everything that's come before has this variant been seen before if it hasn't been seen before can we infer things about it based on where it is or or other things that we know so um I mean there's a few things that would that would help one um and and this is an ongoing Challenge and it's a moving Target over time we've we realize and and we when I say we not the company but the community realizes that uh there might be additional pieces of information about a given observation that are important for us to be able to take full advantage of it but of course you know if that happens in some year the years you know before that you know when people were collecting all that data might be missing those pieces um so there is a certain I think it'll be hard to make a a single unified sort of data format even if you which would be great right I mean so one if we had a public database of all of the relevant observations um that would be a huge step in the right direction but of course there's a challenge in creating such a thing because we don't actually know all of the details what are the kinds of detail what are the fields I don't know the kinds of details that we want to we want to capture and that's probably going to change over time so you need a system that allows for that flexibility and you need approaches that are robust to those kinds of variations like all right so some data has some detail and some doesn't but we can we still need to be able to do something with it yes uh so there's there's those kinds of things I think that um it's interesting because in a group like ours uh you know we have this overlap of expertise between the specific uh genomic and and biological domain knowledge and an understanding of data reduction and data organization but we're not uh you know world-class experts on either side right so um it would certainly be interesting uh for there to be communication between groups like ours that are so and there are other groups that are working developing biological ontologies or trying to do things like this and people who are um really like 100 data scientists may be thinking about some of these problems um in a more focused way um and um and sort of have that conversation because yes I think that there probably are um pieces that we could break off of these problems and have the community sort of turning on them and I think that happens to some extent but um there could certainly be more yeah so that's I was really motivated by observing all these groups working towards the same goal but maybe you know as software developers uh we can actually facilitate this through open source and it actually strives are working on it's probably a very good way to bring this knowledge together right so we probably need good public anthologies as well because if this is something where you're hanging all these pieces of knowledge onto right so this is kind of a backbone and you know and so this is something which you know computer science understand right so you know if we have a good public Anthology which we all can agree on that probably will move at least you know a long way so so now I'm really interested in how to make that happen yeah so yeah I would say that there are um there are attempts at that out there um one of the sort of fundamental challenges and maybe I'll talk about it a little bit later as well of building a good biology is that it turns out to be fairly labor intensive and you can certainly with software develop techniques uh to facilitate maintaining uh quality and especially as you you know increase the domain and increase the coverage um but um the you know the what we find are the ontologies that are public just don't have that amount of attention being or you know resources being devoted to them um either to uh sort of maintain the quality or even to maybe even to test the quality right and so at least if we can get to the testing part we can say okay here's some good techniques for testing yes then there could be you know you know we we would know first of all we'd have a metric of the quality right we would know what the quality was in any given perspective and we would have and people would have things to work on they wanted to come and contribute to a project like that it was like well here are areas that we know the tests fail in the following ways so I think that would be a huge step in the right direction this I think this sounds like something which software Engineers can really relate to through test driven development like if first startable writing tests which is your metric of quality rather than you start kind of populating and measuring so that's that strikes me as something issue as a community can probably try but you know thank you very much Jeff it's really interesting and we're looking forward to your talk thank you thanks