Devreal

sftext.org: Alexy Khrabrov interviews Ilya Kreymer about Internet Archiving

sftext.org: Alexy Khrabrov interviews Ilya Kreymer about Internet Archiving

Recording: sftext.org: Alexy Khrabrov interviews Ilya Kreymer about Internet Archiving

hello everybody I'm Alexi crabber of the organizer of SF text meet up and hear our location with Ilya Kramer our speaker today who talked about wayback machine common crawl old web and his basically master for having the way the web looked at some point in time and we're going to talk with Ilya about why is this interesting Chloe up it's great to have you here great to be here thank you for having me speak today so so fascinating talk one question is why is this interesting and why is this interesting to you well it is interesting to me and I think in general because so much so much of what happens in our in our lives happens on the web hmm and things on the web disappear people don't realize that it may it may seem that that the web is permanent but really it isn't and URLs and content on the web disappears all the time I don't have stats on me but in a slide I had a other informaton how it so it's sort of part of history it's I think that it's an incredibly important thing to do and what i'm doing i think is building new technology to try to keep up with with preserving the web as the web evolved so a lot of the tools that have been built to preserve the web have know many years ago and the web is constantly changing at a really fast pace and so new tools need to be built to in order to do keep up with with the changes in the web and at the same time you know Charcot has done a great job of saving the early web and I'm also interested in kind of making sure that that's preserved using old browsers and that we can sort of remember how the web was and then study it and and refer back to it using the environments as close as we can to the environments that existed in ok and so maybe you can explain to folks who are new to this as you mentioned multiple things which are horizontal under the graph is a company's nonprofit conquer all is nonprofit we're back machine is is a tool which internet archive was using so can we kind of explain we should these things are kind of different entities and how historically they appeared one after the other and kind of what this field looking like today well and yes and I think that the project I'm also working with a nine nonprofit called rhizome dot org and then they're a non-profit committed to art preservation or digital art preservation and in German current preservation and so Weber quarter is also a nonprofit project and I think that's probably not a coincidence perhaps because I think that perhaps preservation of its well it'd be great if it was a perhaps a for-profit commercial entity i think that the nonprofit aspect of it makes it more durable in a way and more more perhaps more difficult but also more durable in that will see that this is an organization that is trying to preserve things for the longer term rather than make a new product and you know make a bunch of money in the short term with for investors for example and so I think that the nonprofit approach definitely makes sense for longer term digital preservation and then perhaps is the best approach I don't know if yeah then that that's sort of my feeling about it right now is that nonprofits are a great option for for for software development also I and in this in this field in particular mhm and I guess my question is also why do new organizations appear because the original profit so it would be logical for them to all unite but you know we see Internet Archive then the Sekhem on crawl right let me see this new generation so what kind of as kind of the creation of new archival organizations well I think everyone's you know I think it's good that there's multiple organizations that are I don't think that there should be just one archive that that's dangerous you know what if that archive but what if that archive isn't is inaccessible so I think the more archiving organizations there are the better enter archives been around for a long time they've been crawling the web and also other other content that they have an archive of many different things available to to the public some things and then common crawl for example is more focused i think on perhaps more more for researchers in terms of crawling data not necessarily full web sites but crawling data and making it publicly available for researchers hmmm and what I'm trying to do with web reporter is kind of distribute all of this so that rather than have one one big archive I would like to build tools that enable many smaller archives and that perhaps different than than than what the entire archives mission is in some sense so yes I think it's really excitable do you describe because we're recorder lets individuals see if that their own histories like and I think this is the only feasible way to solve this dynamic content problem right so because I think a few years ago there was this question was behind the ? right because there are the parameters which can figure a URL to basically be an API to a database you cannot crawl the database and so now what's happening basically everything you know Facebook is just giant database and you cannot crawl it and and so and most JavaScript apps become single page apps which basically produce content based on very clicking so it almost sounds like you cannot really have the web but you can archive some snapshots right and that'll be kind of discretized version of the web right right that's definitely yeah that that's that's a fundamental problem is that you only see the client-side not not the server side however there is a lot of stuff happening on the client side so as as more client-side applications become more sophisticated you know right now focusing on one specifically sort of archiving HTTP traffic but you could also archive JavaScript events and then figure out ways to archive the client-side more effectively as well although it is true that that and sometimes you're still creating snapshots as far as the the server status is is concerned I think that there's more more opportunity to to archive the client side as well mm-hmm so what do you think about multimedia contents obviously a lot of these websites and radio stations and TV stations and so are you trying to capture the sound and video or is it up to the sites like YouTube which can actually host them save the videos what do we do about this mom text data well I mean the idea with with web recorders that it allows users to archive anything that that's you know what what their personal browsing experience know is so it was owned including set the sound and video so if you are browsing a YouTube video it should be able to yeah I mean it should be able to our archive that and that's that's up to that particular user as as with anything Allison in the in their browser and of course you know YouTube is certainly not should not be regarded as an archive in any way because YouTube videos get removed all the time right for a variety of reasons some valid some less valid and so if a user needs to have an archive I mean the user could be a museum for example or more or a library that I would like to have an important video that was uploaded to YouTube now if YouTube besides that that they take it down then that video is lost to everyone including a memory institution that that should have a copy of it and so with with tools like Weber quarter it gives institutions and individuals the ability to archive what they what they see is important to them so basically you say that I'm owning my information consumption experience and because you know I you know it was a time in my life when I was looking this V at this video I all the time and have a record of this regardless of what you know external world is doing the source of this video I kind of in my private space I should be able to enjoy basically right like replay of that so this is their interest of cop caption personal time so I'm just you know kind of just good to go back like it like it took a little bit about your career right because it's not a very common thing to join a non-profit or having the web in Silicon Valley weather is like you know million start-up opportunities so can talk a little bit about you know how I got involved with computers the internet on kind of how did you arrive at this occupation well I've been a long time ago I worked in video games that that was actually the first field i started in and I worked in video games for a number of years and I still have many friends in that industry and that's a then I decided to take a break from from some programming and I I didn't actually focused on the focus on playing accordion for a number of years oh I took a break and but then I felt I kind of decided that it was you know my savings right ready at and that I was time to to look for to do something that I found to be meaningful and I had a friend who ever she wasn't here tonight but a list a good friend of mine who I played music without who works at the entire archive and he told me that they were hiring and psycho well I've always known about the Wayback Machine I thought that there was a really awesome tool wouldn't be cool to work on the way back machine and until I applied and that's how i started working there in in 2011 this is a serendipity yes and then meeting the right people yes so it's sort of a combination yeah it's an interesting stuff I can't help asking you about the accordion because I entertained the idea like I i I'd like to eat the idea of knowing how to play the accordion how hard it's actually learned playing the corner uh well with with practice it becomes easier like with anything else does it play a little bit of the piano yeah if you play the piano that makes it a lot easier but yeah I mean it that that helped me definitely with the with the right hand it's pretty much about the same in the left hand is I think a bit simpler even in the piano so I although lately i've been i've been so busy with web record i I'm playing as much so I need to get need to find another balance you know and play some more interesting yeah so we know we kind of try to combine art with data science that our dated by the by events so we know it's really great to see musicians kind of a data scientist kind of the same person so we like who always would be well you know happy to see you play accordion that you know and talk give it a can you talk like that that's a fantastic combination oh thank you I guess thank you again for for having me here yeah I think it's a yeah and I'm happy to have been able to present here those hugely entertaining and very interesting I think like a lot of you know all the data scientists and kind of all the kind of text miners they need this data right like we have this record now we can go play with it so I think that's really give us a lot of his eye so thanks again and we will you know happily curvy all in all talk in the future great well thank you thanks yeah yeah thank you