data.bythebay.io: Greg Lindahl, Data & Metadata at the Internet Archive
Recording: data.bythebay.io: Greg Lindahl, Data & Metadata at the Internet Archive
I'm actually especially pleased to be presenting in the room named after Martin Gardner who is a famed computer scientist and he wrote a column for the Scientific American for many many years and when I was a kid I loved that column a lot so so a little bit about my background so previously i joined the internet archive five months ago full time and before that i was a founder and CTO of a search engine block o people always want to know so we know what was your thing well we're like google only less successful was it what it turned out to be in the end but if you see my eyes light up about sort of web scale data and web scale search that's the the reason because i've been doing it for a long time now so the internet archive probably a lot of you are somewhat familiar with what we do so we're actually a nonprofit library and our very modest goal is universal access to all knowledge and we're not we're not kidding about that so we have extensive collections of a lot of different stuff so we have a bunch of software titles and these are actually old you know pc games or Apple two games are extremely popular part of our collections but we have you know we have videos and we have physical books and we have audio recordings and we have television including an archive that's very popular right now so we record TV and all the important primary TV markets and we analyze both the newscasts and political ads who's running ads for and against whom and who's getting uh basically free press which Donald Trump has very effectively exploited in the early primary season and so if you see articles that are written about you know thus and such as being attacked by this super pack etc said are very often the underlying data for that is actually provided by us in our TV archive we have a bunch of e-books that we don't have physical books for in addition and then finally we've got the wayback machine which is a web archive and it has a total of four hundred eighty two billion web page captures from 1996 to present and so this is you know if you you know get a 404 somewhere for a page and it was popular enough it's likely that we have it in the Wayback Machine so the total of all this data adds up to 25 petabytes of which about one-half is the the Wayback Machine web archive the second biggest collection is actually a video and things like books for which we store both the images and the texts are much smaller and cheaper for us to store so this collection is growing rapidly over time so for those of you who are big data fans we've got both variety and velocity going on so as you can see if we've been growing exponentially for quite a while the web is no longer growing exponentially especially the amount of actual content on the web definitely isn't growing exponentially however you want to define that by any measure but there's certainly a lot more web spam every day and we're adding two billion web captures per week at this point so if you want to talk big data we've we've got a lot of different V's but especially velocity and variety so if you've used the Wayback Machine you're probably familiar with this page so you can paste in an earl and and find out what that page was about then you go to this page which is not very illuminating along the top you can see sort of all the different times we captured it and these little blue bubbles on the bottom they get bigger as we've captured it more times in a day this doesn't really answer the question that you probably want to know which was wind this page significantly change right lenox calm it started it was owned by lenox international and then they got transferred to the Linux Foundation etc etc there's been lots of changes in this page over time and so one thing we've been working on at the Internet Archive is better exploration tools to surface what we actually have and so this is an example of a tool that we're going to launch with a different UI in a couple of days along with the rudimentary ability to search the Wayback Machine and as you can see I've got it divided up and a major and minor changes over time which for lonex com alas it changes constantly but for a web page that's a bit sparser it this has a lot of interesting information to it another thing you can see people always imagined oh yeah the Wayback Machine it goes out and it crawls stuff but actually we didn't call anything on our own really until 2010 so from 1996 almost all the data was provided to us by Alexa internet which is a commercial company that was co-founded with the Internet Archive they were twins at birth and Alexa still gives us all the information that they crawl over time which is based on popularity from the Alexa toolbar we have 400 over 400 libraries that work with us for crawling so especially the national libraries for various countries they've been tasked with saving the archiving the web for their country and generally those folks work with us to do those crawls and almost all of that crawling is also visible in the Wayback Machine so here with this mouse over you can see all the different people who've worked together to make captures of lenox com over time so the Wayback Machine has a lot of structure to its crawls and that wasn't visible until in a few days so what do we do well the archive so we're really great at not losing your data so we have a homegrown system that stores sort of everything in four places we have a philosophy called that I'd like to call metadata everywhere where basically wherever the data is there's a copy of the metadata and so if you find a disk drive that may or may not have good information on it if you plug it in you can look at the metadata stored next to the data you can be sure that the data itself has the right check sums and there's also enough information that you can go and figure out if that's the most recent version of the files another thing that we're really good at doing is making derived files so whenever you somebody uploads a video it is usually not in the format that you want to present outside people either it's too high resolution or a wacky format or what and we're very good at transcoding and OCR Aang texts and generating epubs for book readers tablets and all that sort of stuff so so what's what's easy what's hard to us so we're extremely careful with the data and when you upload files we make a big differentiation in between the fundamental files you uploaded and then the derived file so we made from it and where we try very hard not to lose the the original files and we're a little bit more relaxed about the drive ones because we know we can always remake them the amount of compute we have / storage is dictated basically by the server design that we use so we have a for you server brick that we use for almost everything and it has two cpu sockets and 36 disk slots and we put a terabyte Seagate shingle drives in those so we have more than 5,000 of those shingle drives out of I think what 20,000 total disks and most of that compute that we have available which is quite limited it's dedicated to the drive system that i mentioned before so it's out there transcoding and OCR eng and doing all this stuff that needs to be done in order to make stuff available now because of the way the drive system is built to be very careful about not damaging the original data it doesn't have direct access to the fundamental storage so what happens is is when a drive job starts it gets a copy of all those fundamental files it builds some additional files and then later those additional files are copied back into the primary and secondary storage so this does something which is the reverse of what you're used to have happening that the efficient way to process a lot of data is to move to compute to where the data is and instead we move the data to where the compute is and so that is something which is important to us because we're trying very hard not to lose the data but it makes us very sad when it comes to computing in turn of the structure so let me give an example of an easy thing which is made hard by our infrastructure and this I'm not kidding about this somebody came up to one of our library folks in Iceland at the Internet Archive preservation Council and wanted to know how much Luxembourg text there is in the web now I don't know about you but I had no idea that Luxembourgish was a language but it is and there's actually a language recognizer that works for it so the chromium folks at Google have a CLD to dash full library that recognizes not only all of the national languages of Europe like luxembourgish but they also recognize a lot of the regional dialects of which you can either considered separate languages or not it's like Catalan and Manx and all sorts of stuff like that question so so Luxembourgish is written differently if it wasn't we wouldn't be able to play this trick but indeed but I'm very aware of the of the german thing because i had a bio ish girlfriend for a while and indeed yeah the the u.s. kind of you know has a similar thing and there's a few words that you can spot that are different for different regions of the u.s. and then the pronunciation is very a little bit but oh my god gota to Britain if you want to see wildly different accents in English Britain is an amazing place that you frequently have to concentrate really hard we're in the bar trying to have a conversation with that guy oh my god what's he speaking oh it seems to be English anyway so so we have this library and we've got 12 petabytes of compressed texts that we want to run it over and and then we have to store it somewhere in our system and we're not really designed to store data like this we we kind of have the ability to store metadata but it's really kind of limited so some days my job is like this yeah big data I hate it and so so your first big takeaway is if you work for a start-up then you get a chance and if you're on a greenfield project for an established company also get a chance to to pick your tech stack and then you can also change your mind if you do it early enough and the internet archive is a fast moving organization but we are 20 years old and we have requirements like not losing the original data and so at last we don't really have the luxury to change things too much okay so onto the Wayback Machine so the wayback machine plays back the web and so this is when you want to go look at something and today the way you play back a web page is you paste in an exact Earl and so if you've got a 404 somewhere on the web that's that's sort of an easy example and and by the way we're working with browsers to get integrated directly into that so when a browser whenever it sees a 404 anywhere on the web it's going to come ask us if we have it and put down and if we do it'll have a slide down thing that will explain that that this link is archived at the internet archive a little bit about the internet archive is and if I think you sort of go to it so this is going to be visible to a lot more people so the basic index right I said we have 482 billion web captures and essentially we have a text file that has 482 billion lines that sorted into sort order so there's a there's an order that you can take an earl and it flips the host name around so you can sort it it's called a cert su RT and so you can then use an ordinary sort program to sort that this huge file and and so of course you could binary search through it right but in order to accelerate that there's an index for the index and well okay so you have to update the index to and so we actually have four different indices for this data we have the all index which contains everything that's older than two weeks we have a couple of intermediate indexes called delta full and delta delta and as you can guess by looking at the names of course there was only one in the beginning and then finally if you hit save page now which is one of our most fun features for reporters or anybody who wants to see something juicy on the web and was to make sure it gets preserved for eternity you can paste an early in and hit save page now and and that Earl can be played back immediately thanks to a Redis database so this is a you know awfully complicated but it it needs to be for a reason right that's a big table so so so the nice thing about this is that it's fast but there's a there's a footnote to that right so you hit save page now it goes into the Redis immediately can be played back it moves through delta delta into delta full into all over two weeks so of course it's more complicated than that well okay so all these indices are actually sharded three hundred ways and you may be asking you at this point you probably wanna raise your hand and ask me why the heck don't you use you know something like you know clustered Redis isn't that capable of doing this and well you know it didn't exist when we invented this thing and so this thing is this particular system is 15 years old and it grew up over time right so it started with only the all index and then delta was created and then it got split into two pieces and it used to live on data cated hardware but now it lives in our main infrastructure so each of these things is charted three hundred ways and it lives on spinning rust over try our architecture so for for better or for worse right when you play back a single page and it includes images and JavaScript and CSS a single playback of that page you know basically has to look up ten different things if there's a page and nine embeds so then I mentioned that we're hoping to integrate this into browsers and that means that suddenly we're going to have we're not sure between five and a hundred X our existing look up right it turns out that a lot of usage of the internet archive is by BOTS currently and so I hope to get rid of the bots and that might get us down to only five x but maybe it really is 100x I don't know so so holy crap so if this is a Greenfield situation right I'd wander around and say you know who is it who has you know some sort of no sequel thing for a single table that's 60 terabytes in size and several people would raise their hands you know yes we already have people like that using our no sequel database in that fashion and that would be awesome but that's not a choice we have necessarily and then there's that that footnote I mentioned well so if there's any hiccup in our infrastructure it severely affects playback because playback is looking up 10 different things in four different indexes which are sharded three hundred ways and so if you want to fetch a book and read it that's fetching one thing but playback fetches a large number of things and so there's sort of two pieces to this the first no plan survives contact with reality and secondly it's never good to be the only person complaining about the infrastructure because it isn't good enough your colleagues look at you funny and then they did tent they want to decide that you're not using the infrastructure right but unfortunately that is the the look up so I'm hoping to actually treat this like a Greenfield thing and go out and and completely replace it but organizationally this is an interesting challenge right because your hindsight is twenty-twenty oh yeah we should have done this 10 years ago five years ago when it became possible to do with more standard technology but it's it's super hard to not discontinue sinking into the quicksand when you have a system like this so that's enough about me complaining about playback so let's build a search engine for the wayback machine because that's what a lot of people would love to do in order to find stuff in the Wayback Machine right you have you had an old website and you can actually remember what its domain name was if it's completely disappeared off the internet then Google won't help you either so in so much as people find things successfully in the Wayback Machine now either they found a link somewhere that sent them to a 404 or a redirect you know or a software for or whatever and it was clear that the data used to be at a particular Earl so they can go ask us about that Earl or they can do a google search and they can discover get a hint of where it used to live but if neither of these are true because the web page disappeared in 1998 and all references to it are basically gone then what you need is a search engine so let's build a very simple search engine so building a search engine for the web is easier than ever and and I say this is a guy who's built one it's sort of peaked at I don't know in difficulty and maybe 2009 and sort of went down so as long as you're able to figure out what's web spam versus what's web content and that for us that decision was made a long time ago so we have we have what we have so so all you have to do to build a search engines super straightforward you iterate over all the HTML files and you extract the outgoing links and the anchor text which is sort of the underlined thing that you click on because that is actually the number one thing for characterizing a web page and you invert those outcoming links and incoming links and then you throw those docs into elasticsearch and okay so now you have a little sub iteration where you often you torture elasticsearch to use ranking information that is not anything like tf-idf question so so indeed and and the challenging things so I've I built this search engine right right right now I'm only building a prototype because if I do this right then I have to show you a temporal answer right so if you had a website that was very popular for certain terms in 1998 but then died out and some other website came along in 2002 and was popular for those terms until the present then the results need to show both of those things and so I'm not explaining how to do that because I don't know how to do that yet I have some ideas but I have yet to test any of them so indeed that is for the wayback machine in particular you know all other web scale search engines only want to show you the present and we want to show you the past but for this very simple thing right you just throw the the docs into elasticsearch and let's see so okay so we're going to build a bad search engine so remember I mentioned that we have limited compute available in our infrastructure and this whole copying thing so we iterate over all the HTML files that takes 100 days in our infrastructure well that's a lot of data right and then you invert the outgoing links into incoming links so we do have a little bit of help in our environment so so try number one so so in order to avoid all the extra copying because it will take a lot more than 100 days otherwise one of my coworkers who's a very smart guy ordinarily wrote a thing he tested it on a small fraction of the web and he's like oh I have to make sure that the the thing that extracts the the links isn't going to blow up right in memory usage um so I'll limit the input data file to four megabytes it said the HTML page 2 4 megabytes and he tried it on a few million web pages and it worked great and then he started it up on friday afternoon and left for the weekend so how many of you can guess what happened yeah so there are of course a pathological web pages out there that even with a 4 megabyte input file can generate in many gigabytes of memory usage in an HTML parser and and you can probably imagine you know how that's done right very straightforward thing and so not only do you run into people who've intentionally generated such HTML files because they think it's funny and it is there's also people who accidentally create files like that because they have no idea what they're doing and on the web you will always discover both of these effects going on that day you're being you're running into folks who who thought it would be funny and and they're very smart and they are and then the people who just did whatever and and oh my god it's a mess so that was the so it only caused a 15-minute side outage you know and thank God we're a non-profit and nobody has super high expectations for our us being up the second half of the problem was is that the limited scope of our Hadoop cluster meant that we only could invert and index the root of websites so alas for example the BBC food archive which is about to drop off the internet because the British Tories are having their problem you know personal problems and they want to go delete half of the BBC's old content it's a it's a subpage on the BBC page so if you if I wanted to build a search engine in which you could type in BBC recipes or BBC food and end up at BBC co uk / food / alas we weren't able to generate that but we generated that thing that kind of works so this is the again the this is not the UI that's going to be launched in a few days but so you type in I was at the scale conference when I made this slide so you type in dot and it comes up with lots of things which are not scalable cuz of course that's not a particularly popular domain so clear dotster net blog engine etc etc so this is a bad search engine for the for the the web so this isn't like Google and and this is you notice that I've had searched with a line drawn through it and all the previous slides the problem is is that when you put up a search box on the internet everyone expects it to work exactly like Google and so no matter what I tell them they'll type some words into the search box that describe the website so i'll type in BBC recipes quinoa no matter what I tell them to do because they expect that to work because it works in google but this is as much stupider this is not only search but but nonetheless ok so we ship something that kind of works but the challenge in our infrastructure is we do not know how to scale this 24 and 82 billion web page and if I want to put this into perspective it appears that the primary index at Google is about I don't know 15 billion web pages and the Supplemental index this is why they used to divide it up it's like 85 billion web pages and it doesn't really need to grow because there's nothing any good web pages on the internet right so this is a very difficult problem and then part of the reason we have 482 billion it's not only because we've crawled crap that probably google would be smart enough not to crawl but in addition we have mini mini captures so for example delonix com homepage we've got what is it it was some 100,000 captures or some very large number that have to somehow be combined when we go through and and and and index it so that's that's a little bit about web search so so this is a conference and about data so i should mention by the way we have an API so the the in the beginning i gave that list of all the stuff we had right so we've got we've got books we've got video we've got audio and majority of that stuff is stored in items where an item is a book or a CD with a bunch of tracks in it etc etc and we have an overall search engine for the website that that finds those items and we've made all this available via a python tool that basically covers up for the fact that we don't have a good api so the python tool basically puts a bunch of spackle over that and then paints a beautiful mural and so pip install internet archive and you will get both command line tools and also some beautiful-looking python classes and eventually our actual api's will it be as beautiful as this but for now that's the way you can get it stuff and anything that's available on our website is actually also available through these tools so if if you're interested in either mass download or mass upload then this is definitely the tool for you so why would you be interested in mass upload so I run into all kinds of interesting people in the Silicon Valley and so I ran into this guy a couple years ago and he's like you know I you to work at xerox parc and I'm a hoarder so many people are I have the stack of memos I'm like oh my god scan all the memos and so he's made a special collection with us and basically he's going through one by one and asking all the people who are authors on those memos if he can make these memos available in public so at the very least we have an archive where the memos can be the metadata about the memos can be seen but if he's cleared it with all the authors for a particular piece of content the content it's also visible on the Internet Archive to everyone and so that's one of many many reasons to want to be able to mass upload stuff mass download we have a lot of interesting stuff and so for example I mentioned our video game collection which is extremely popular but we have a bunch of other interesting specialist collections and I don't know what hobbies all of you have but we probably have something special for your particular hobby somewhere on the internet archive and so you can start by surfing around the website but when she sort of narrowed in on a collection you may find that having these command line tools make it a lot easier for you to deal with them and I have a few examples of using the command line tools here so this is basically there's all of our items have names and so in this example trip down 1905 is the name of this particular item and I suppose it's one of those Victorian travel books or something tripped down the Nile I don't even know and so it has a bunch of different stuff and so as you can see we're showing you know downloads for video or you can download the same video in several different formats so I mentioned transcoding so we make mpeg-4 is patent encumbered and OGG video is not and so we make there are many places where alas you have to use impact for because that's built into the hardware and otherwise your tablet will run out of battery halfway through the movie and that sucks but we also make things available in free formats like video and you can look at the raw metadata and all that kind of stuff question so so we want equivalents such that one hundred percent of the information on the web sites also available through these sorts of interfaces and then four people actually using this if you are one at a time then you will not overwhelm us we have a lot of bandwidth to the Internet and you will probably be limited by your local ISP and not by us if you're in an ordinary sort of situation so no we don't mind people downloading a lot of stuff from us just please God whatever you do don't start a hundred downloads at the same time that will make us irate so question so each individual item has licenses associated with it and and there's a wide variety of different stuff on the internet archive and for anything that's been uploaded by a third party we have generally not verified the license in any way and and there's some stuff that is up on the archive and is actually restricted as to who can see it so let me give you an example of the variety of licenses in our book collections so we have a public domain thing anybody can upload into and so you know corey dr. Rose books which are available on a very generous Creative Commons license are in that collection there's some stuff in there probably actually somebody falsely said that it's available to anyone when we get dmca takedowns for those things and we take them down so there's also the books we scan or our library partners scan and these fall into several different buckets there's things from before 19-23 we call this pre modern books and so in the US copyright laws very straightforward and all those books are in the public domain and so we make those books available to anyone if a book is younger than that it sort of falls into two different categories one category is that we do not own a physical copy of the book and that book if we scan it is a part of our print disabled collection so there's a special exception in the law for libraries nonprofit libraries to be able to scan and make those books available to print disabled people so the people who are hard of seeing or are legally blind there's a process they go through to get a particular key and that special key you know so their doctor talks to organization related to the blind and they get a particular key and so as a result they're able to read those books at the Internet Archive and we have eight hundred thousand books that are in that category and then four books that are modern where we do on a physical copy there's 250,000 books in that category so we scanned the books and we put the physical book into a shipping container in our warehouse in Richmond California and for each physical copy of the book that we own we will loan one electronic copy at a time and we hope that that library will grow from 250,000 books to perhaps 10 million books over the next five to seven years and that will us that will give us the ability to be a library for the entire world and have a very large number of the important books available to folks around the globe so of course we'll end up not owning enough copies of some books and other books will never be checked out so be it but that gives you an idea of the variety of licenses and this is only talking books it's even more complicated for music and many other different things you would be surprised that we have a lot of for example modern video that's people's home movies and things that are available on under very generous licenses so so don't really assume that you know since almost all video is post 1923 that we have very little accessible video from those eras it's it's a much larger than you think surprisingly question Wikimedia so we work with the wikimedia foundation and as you know their database downloads provide the complete edit history for every single article and so while we are ordinary web archiver archives those pages every once in a while but Wikimedia has a better archive of their own data thanks to this edit by edit phosphate the places that we work with them that's that improves what they have one area is actually in transcoding so Wikimedia has a lot of video uploaded and then so much is available in multiple multiple formats we actually did the trans cutting for them and that was interesting because that was a project we're sort of engineers worked with engineers to make that happen without either sides bosses being available that was going on so that was pleasant suppose it surprised to discover that was going on Wow who knew that we were helping each other out that way secondly for external links and wiki pedia articles there's sort of two needs for that so first off you want to be able to see you know the content at all that was on those pages and as you know right you cite a web page from from Wikipedia and it you know twenty percent of the time it disappears within the first year and but secondly you would also actually liked an archive of that web page on the exact date that it was cited right so a citation for web pages an earl and a date and not surprisingly the name of a web archive in the wayback machine is the Earl and the date it was captured by us and so we've been running a bot for a long time that's been watching all the edits to all English Wikipedia pages finding the external links and immediately archiving them and then there's a bot called cyber bought two that's been going around testing all the external links and if it's in the archive it'll add a link to the link in the archive or if it's not available on the web it'll flip and it will make the archive link available as the main link and then it'll say was at and the original link which is now dead so that is mainly for the English Wikipedia Wikipedia as you know is very editorial and they had a vote I guess six months ago with the editors picking what the most important new initiatives for improving Wikipedia and then the thing that got the most votes was strengthening this bond with the internet archive so we're currently making a crawl of 123 million external links from all 608 wikis that were part of Wikipedia so that we can do that for not just English and not just recent but sort of for all time across the entire wikipedia collection no matter how old the articles are so we won't be able to get all the data because some of it disappeared some of the web pages disappeared a long time ago but we have a pretty good coverage of what they got and we're very happy to be working with them because they do what they do very well and and you know we do some things that can really enhance what's going on on their sites and vice versa so I have a few comments to go so one interesting thing about the Internet Archive is pretty much any project that we do turns big at the drop of a hat so I described how we have an existing search engine that works over the items in the internet archive where an item is a book or a CD and that's a hundred gigabytes of index but imagine if we wanted to extend that system so let's say we want to index all the individual tracks for example on a CD well that 5x is the size of that index let's say we want to go in for books and make algorithmic metadata so you know normally a book you index only the card catalog card so basically literally the information that would be typed on that little you know three by five card and that's the state of the art today and libraries but we know looks like anybody how to go and through a book you know and use in the extraction to find out a mention of all the people who are places and things that are in a book and it's sort of a score every round Horton those things were for that book and we'd love to index that and that's at least five Xing the index and then we've got information that's locked away and the wayback machine so for example there's a very large number of PDFs on in the web and to be really nice to be able to index the titles and authors or at least the front page or maybe the abstract so if it's a scientific paper you'd be able to actually find that in the web in the wayback machine in a way that would be much more sophisticated than the very simple search engine I showed you that we built and so suddenly we went from 100 gigabyte index which is oh that that's so cute it's a little tiny thing to 12 terabytes of index and oh my god so a few more takeaways so organizations are made of people and so technical debt isn't actually necessarily about technology so I would love to go and replace the drive system that we have with Apache my size but the drive system has been worked on for 20 years and it has a bunch of specific things that deal with certain situations this particular uploader uploads this stuff which which gives us indigestion and so drive has a workaround that makes that not cause a problem with our system so there's 20 years of heated hacks in that system and there's also a bunch of people at the archive some technical some not who know how to operate this drive system very effectively and so it's not really about the technology it's about the people and finally I'd like to show you one more toy since i have a few extra minutes so i mentioned the very beginning that i loved the data that we have and so not only do I love the wayback machine but also we've got a large number of books especially the modern books in libraries right the pre 1923 books not that many of them get used heavily almost all the usages of post 1919-20 three books and so what I did for this this this was a volunteer project I did with the archive before I went to work for them full time and it's a it's a visualization of information in book so I took the 80,000 at the 800,000 books for which we have stands available and I took all the sentences from those books and I took all the sentences which had a day and I extracted the entities and those sentences and then I built this thing so what it does is as you type in an entity here and and I cheated I used wikipedia article titles and they're synonyms for entities so Joan of Arc it's a person and she was you know very popular in the 15th century in France and so the histograms are the number of mentions with a date in the same sentence and so the first peak you see there the tall one is from when she was born until when she died with the tall peak being when she led an army that relieved the siege of early on and 14 29 and if you click on that that the bar on the graph it actually shows you ten sentences from ten different books that have that date and Joan of Arc in the senate's and so this is a way of exploring the content of books which is not a search engine but it is a search engine and then if you're wondering what that all the stuff on the far right is now 1902 2000 basically that's modern scholarship that mentioned is Joan of Arc so this is a book vigil visualization engine and so I'd like to finish off with this slide so if you're interested in looking at any of the things that I've shown all the the stuff i've shown the the the search engine and the way back explorer are both currently available but we're going to within few days haha we've been saying that for a couple weeks that basic search engine is going to be available and then the date visualization i showed you is it book start archive labs org and then finally our Python code for accessing our metadata so thank you very much I have a few question number one can you share your use of our elastic search how big is your cluster and did you run into any gacho number two who can access the content and use it for you to first accessible by everyone and we had to deal with the conga rights a related issue did you consider sir question did you can see the integrate with the copyright management tools because apparent you do have to deal with that so let me ask the second one first so I am I am I used to be I used to pretend I was copyright expert but since I started working for the Internet Archive I no longer give opinions about copyrights however the internet archive works very closely with the eff the Electronic Frontier Foundation and pro bono lawyers that help us out and the policies that apply that I discussed earlier to various different kinds of items so books of different ages and a differentiation for whether or not we own a physical copy of the book or not we're all reviewed by those lawyers and received a reasonable blessing from them you know remember you know libraries and archives are not like other organizations and and there are special things that apply to us for example the ability to scan books to make them available to the print disabled is a specific thing for nonprofit libraries so that was the second half of your question for the technical first half so the cute little 100 gigabyte index actually runs on a cluster of you know 10 to saket knows each of which has one SSD and unfortunately we did not do a great job of capacity planning because we don't have the ability and that cluster to increase the size to 12 terabytes like I'd like to because it can be challenging you know you're going and you're looking for music in the internet archive and if you remember that the song name but not the album name you can't directly find it in our search and that's that makes me sad I'd love to make you know all this stuff more available on the web the way back search side we don't have a permanent home for that and so basically the index of only the home pages of websites fills the little for machine cluster i'm using for that but that's just a toy and i hope to make that one much larger over time but it remains to be seen how deep i can make that index because so blech oh you know built web-scale index with only a hundred billion pages and if i built an index that was larger sorry i was only 1 billion pages and if i could build a larger index since it doesn't changes rapidly but there's that temporal thing blah blah blah so blekko spent 63 million dollars and that's you know four times the yearly budget of the internet archive and I don't get to spend all that money so it's definitely a resource-limited thing for web scale search and it's unclear to me so far what the compromise to fit within a reasonable cost structure is going to be I don't know all right let's say we're out of time you