Text By the Bay 2015: Oren Schaedel, Knowledge Maps for Content Discovery
Recording: Text By the Bay 2015: Oren Schaedel, Knowledge Maps for Content Discovery
thank you thanks for sticking around and thanks for inviting me to give a talk so i'm a data scientist at versailles and versa we are an online education platform so we help teachers and content providers put together content so they can teach engage and give assessments sorry and get assessments from uh other students or employees or whoever wants to consume content and what i want to talk to you about today is how we envision content how we curate it how we create it and how we explore it and especially for the exploration part i want to talk about knowledge maps and how we create a knowledge map using taxonomies and how we balance between a generation of a taxonomy and visual use of a taxonomy so it's both informative but also useful okay so first of all content creation in the classroom so when we speak to teachers what we see that teachers are using now in classrooms is predominantly what we all know powerpoint word a few google products but it doesn't really complete the cycle okay so when a teacher wants to teach something they have to put together content so it's usually cutting and pasting content from around the web if it's video if it's text if it's a few images okay they have to distribute it to their students so i have to send out a powerpoint or present it and then when you want to get an assessment back from a student you have to submit the homework okay so there are a bunch of online platforms that do that but um it's still kind of a broken process the frustrating part of it is that we're now with the current state of the web we have some very beautiful websites is anyone familiar with uh snowfall a story from the new york times adventure in up in the mountains beautiful media a lot of different types of media blended together beautiful visualization anyone read the 538 new york times blog now anyway very beautiful data visualizations so the whole idea about it is that we have a lot of really nice content a lot of nice visualizations but is it really for everyone but it's really for everyone it has like a media team a lot of good developers but we want to be able to give that access to everyone else so at verso what we're doing is we're creating tools to both curate and publish very beautiful visualizations and tools very quickly and the idea behind it is we start off with very small components with gadgets so a gadget can be a piece of text video audio drag and drop tools we group these gadgets together to create courses teachers can publish them and share them okay so very briefly so we have a course about photosynthesis we can take videos mesh them up with text add some google drive documents quizlet all sorts of third parties and even go very deep and very complicated like making a time series of all sorts of events around the civil war annotate them with content that's very easy you can just drag and drop it okay so why am i telling you all of this in this context we have a lot of content we deal with a lot of content we want to help people organize all of this content okay so you can take this content regardless of reversal and you can stick it into your own website a gadget is a is an iframe so you can put into your into your blog into your wordpress whatever so we deal with content all the time so the question arises is how can we give people access to a lot more of this information and how can it be very useful for everyone okay and how do we present it in a way that's not too messy both in a way that you can see a very large picture of all the content you want to look at like a 20 000 foot view and also like down to the very small details of the specific content that you want to capture so we're working on is called reversal knowledge map and think try to think about it as a google map for content let me demonstrate a little bit more okay so the objective of the knowledge map is to allow people to both explore large amounts of content exploit it use it and extrapolate it and then share it okay so let me try see if we can get the knowledge map to work with a live demo i'm gambling here live demos have a tendency to do these things this visible okay so this is a knowledge map of chemistry so what we've wow this is all right i'll go back to the let's show it's not really working the way i wanted to but the idea of the knowledge map is we can take a topic like chemistry or biology computer science we can break it down into a taxonomy so if we look at chemistry we started from a very broad view of sorry this is computer computer science we can take a look at artificial intelligence computer science the very the broad topics within within computer science and then with each one of these topics we can drill down to it so clicking on each one of these will drill down and expand more topics okay let me see if we can i'll give a short demo of how we drill down so for example um organic chemistry i can drill down into organic chemistry and get a breakdown into more precise things like organic compounds or or heterocyclic compounds so i can drill down into each one and in this case what we've done is we've taken just the summary of wikipedia pages in each node and display them as the usable bit of content okay now i can use this for i can create a reading list for myself i can add it and email this to myself or i can look at a larger view like for example if i want to see the context of a specific topic over all of the knowledge map let's say like dna i can search for dna and it will highlight dna with contexts of all of the different branches within chemistry so i can read up about all the different ways that dna is mentioned okay so i can summarize that and look at all of the different ways that dna is mentioned in the knowledge map okay so what i want to tell you about is how to how we're thinking about creating this knowledge map in a functional way so you can browse it both to get that nice overview and drill down into the details so what we've used to create the taxonomy is we're using wikipedia as our source so wikipedia comes in a very nice structured way wikipedia provides us with the dump of all the wikipedia pages so it's roughly like 40 something gigabytes in english encompassing around 13 and a half million documents within and that's just the english language okay and we can also get another set of uh relationships which is the wikipedia categories everyone know where the wikipedia categories are yes every page in wikipedia can belong to one or many categories and these are usually human labeled categories lately there are a lot of bots that do that as well but you can group pages into into categories that way it goes into even more detail you can have sub categories and using that table we can create part of what the taxonomy is but when we we started doing it that way we looked at that big sql table and we started creating a taxonomy we very quickly ran into a few problems it grows very wide it grows very deep a lot of the pages and categories that are automatically added on don't make too much sense so we wanted to organize a lot of it but we would also want to organize it in a way that it could fit into a single page and you can see it and you could read it and interact with it okay so besides the parts that we can download from wikipedia like the categories and the pages themselves there are a few very important parts of wikipedia that help us with content curation the first part is the portals has everyone has anyone here looked at wikipedia portals know what they are no i'll show you a snapshot of it in a moment and then there are major topics so each portal has a list of major topics in it and those are the key concepts that the curators of that portal have deemed important for people to to know about those like the highlighted pages of that portal and we use those as a positive control to know that we're making a taxonomy that is comprehensive okay so that once we generate the taxonomy each page of those major topics is in that in the taxonomy or at least a subset or a large subset of them so we know that we have completeness okay so categories for example can you see like the size of the text i don't know if it's clear or not but the categories for example in biology we have a large number of categories sub each category has subcategories some categories do not have subcategories and the topics like i said are just a list of specific pages that are pretty important by the community for that for that um specific subject for biology for example we have um transcription genetics mendelian genetics so these are very important concepts within each subfield okay so when we create the taxonomy oops when created taxonomy jumped around there we start off with the root which is say biology and then we start off building off categories as children categories have children which can be either another category a subcategory or a page in this case we call it a leaf okay and if that leaf is one of the major topics we label it as a topic and we know that we have one of those inclusive sets okay so very quickly this taxonomy grows very big and wide and deep so we want to be able to maintain it and manage it so we can actually browse it so we had a few challenges over here so first of all we wanted to understand how deep we want to go okay we want to go deep enough so we capture all of the topics that we want all the important pages but we didn't want to go too deep otherwise we lose context so for example if we're in biology and we and we mark mufflers or transmissions then we know that we've gone too far and it's just not related to biology anymore okay within the categories themselves some of the categories will have children category child categories or sub-categories that will reference the category itself so we have a lot of feedback that goes on a lot of circularity that goes on so we're going to eliminate that a lot of the nodes can repeat themselves in different contexts so in some cases we want that repetition and in some cases we don't want the repetition so for example what i showed you before we want to be able to capture dna in all of its different subcategorization but we don't want to replicate a node of let's say photosynthesis like we wouldn't want to capture photosynthesis in many different places as to take it out of context okay and finally the constraints around this are that we want to be able to visualize all of this and the main part of the constraint is can we read the labels on the text that can we actually see this as a map so when we look at a map when you look at it far away you see less details you don't see the the names of the roads or anything but when you zoom in you can very clearly see names of the roads but you won't see the name of the city so it's it's all about how do we manage that type of the balance between resolution and and twenty thousand foot view so we came up with uh a few levels of a few rules how to include and exclude different branches so we're calling it the depth versus breadth criteria um what we do over here is we can we take we establish a few rules that are based on the visualization criteria okay so we take a look at how many nodes we have for each level okay so let's say we have a subcategory that has only two or three nodes that is not sufficient for us to make it a category you go in there and it's it's very sparse there's not much to look at okay um if there are however a lot of different nodes in that same category we want to either split them up into other categories that you're not overwhelmed with too much so we cap that in a certain number okay um we also take a look at the number of children of each node so if a category only has a single child we will collapse that uh we'll collapse that node and merge it with a different node or a different branch okay and right and finally what we went with um one of the rules that we used to generate the breadth versus depth is to see that if a if we've gone down a certain path deep enough to encounter one of the major topics we'll stop the expansion of that branch at that point okay if it has reached the maximum level of no of the maximum depth level that we want that taxonomy to have we will we'll see how far away how many levels away are we from the next major topic and if that major topic is not included in the rest of the tree we will surface it up and include it into that subtopic but if it's included in the tree it will stop that expansion over there okay so all of these all of these rules we kind of summarized into three different operations that we're doing on the taxonomy so first of all we remove we can remove leaves um topics subcategories and categories based on some of the some semantic rules that we've devised so basically we look at this at the content of each page and determine its relevancy so what we've seen going through a lot of wikipedia pages is that there are a lot of geographical locations that don't make too much sense to add it so we filter out geographic locations okay another type of item that comes up quite a bit is lists and lists of lists and there's even a page called list of list of lists in wikipedia a very entertaining page okay so we remove we remove kind of those those types of pages okay some of the branches that we some of the other operations we do is we collapse branches so if there isn't enough information to merit a new split inside the taxonomy we'll collapse the branch and try to merge it with another branch that we can enrich data upon so that when you go into a certain node it won't be with only like one page on it or two pages right so we'll have more of a there'll be more richness inside and we'll also split branches into several groups if we have let's say a node with 40 or 50 or 60 children that's just too much content on a page and we'll try to split it in one way or another okay all right so one of the problems that we saw that came up pretty often was text collisions so if we have a lot of nodes on the page we have very large captions or titles and they'll collide with other titles so we're trying to get around that problem so that was that was a kind of interesting problem to get around we started playing around with kerning with different ways to angle the text different ways to size the text automatically based on the number of nodes to select nodes based on the number of based on the size of the text but we did this for a while and then we kind of discovered a very neat principle of working with taxonomies and what we saw is that the deeper we go in the taxonomy the text the titles of the text tends to be longer okay and so the more specific you are the more words you need to describe that concept and we saw this over several different taxonomies so we've gone through computer science biology math and chemistry and this is pretty you and it reoccurs in all of them so we can see that both on the names of the topics and on the names of the individual pages so that kind of root that kind of phenomena it indicates to us that we're going in the in the kind of right direction so we try to minimize the length we try to go for nodes that have smaller text that are higher up in the hierarchy and very long and much longer descriptions in the lower in the lower levels of the text and that it worked pretty well both for both cases and slowly we can eliminate a lot of those collisions it's not perfect at this point but we're still iterating through it all right so bringing it back together so what we're doing with reversal is we're trying to create both immersive experiences with a lot of interactive tools okay we're trying to take a lot of knowledge and and let people share it and bring it in together and finally we're always facing with a challenge right we want to be able to always merge content with very interesting and fun visualizations and that's what we're facing every day those are the challenges that we like handling and also we're hiring to do exactly that so if you're interested let me know and i'm open to answering questions thank you very much i'll put the url for the knowledge map it hasn't been working on the presentation that much but i'll post the url in a moment so it's so what we're trying to it's a balance between the visualization rules and the way that we can automate on generating the taxonomy so when i say it's it's kind of like the an in-between between a heuristic and a cost function so we can't really formalize the cost function that well around the taxonomy so it we are resulting to heuristics and as we're going through different taxonomies we're doing it as a heuristic until we have enough taxonomies that we can generalize it as a cost function but the idea is to generate a cost function around it and to make it a more generative approach and not like a yeah i feel it kind of way so okay so when we have courses so um the gadgets that we use at reversal they're annotated with text okay so what we're doing at the moment is we're looking at matches between the text in the gadget and we're finding the best matches between the wikipedia taxonomy when we have a good match we can use a bunch of methods like tf idf or cosine distance we can match the best uh pages in wikipedia and then from there we can we find the tags or the names of the wikipedia pages that correspond best to that course now we can trace them up through the taxonomy to find the root and from there we can label courses in a comprehensive way that don't that and it depends the way that we're we're doing it this way instead of using like a topic model or a hidden mark or hidden model is that we can control specifically the names of the tags that we want so if we use lda for example we get a vector of labels which may or may not make sense to people okay so lda can come up with a set of tags like beach sandals umbrella watermelon sand banana and that won't necessarily tell you too much about a course about beach volleyball right so we want we want the labels for these gad for these courses to come from a human readable way and we thought wikipedia was a good place to start um we're looking into that we haven't found a good way to do that also on the visual on the visual part because it will require transformation that doesn't have the same level across all fields so for example if you're on if you're looking at so you're three levels deep and you want to find an equivalent uh term for this for a similar topic in a different uh in a different branch it might not go that deep or it might go way deeper so that type of traversal doesn't have for like a better word of trajectory um otherwise it's just the same level that you're traversing so it might not give you the same context does that answer your question all right thank you very much you