Text By the Bay 2015: Jacek Ambroziak, TopicStream
Recording: Text By the Bay 2015: Jacek Ambroziak, TopicStream
foreign about a particular proposal that I that I'm thinking about for innovating in the world of electronic reading ebooks but this is not really limited to ebooks it can be also used with probable education courses specialized material documentation a technical documentation and so on but I did start this work with from from ebooks and it still to some extent ebook centered and it's this the stock is less technical than many of the talks here it's it's really a vision of user interface and how text can be split into into small parts that can can be related um about me I used to be a researcher at Sun Labs long time ago and I worked on something called conceptual indexing which would be really NLP stuff then I also wrote an XML search engine which was used in IntelliJ IDEA and open Office for many years and later I moved to XML Center at Sunland Road the first xslt to Java bytecodes compiler which has become what is known as Zealand compiled and now I'm working for myself and Licensing some of the access LT software and also working on Android applications and experimenting with e-reading so what I'm going to be talking about today is I mean how electronic books function in the 21st century and by electronic books I mean I'm specifically interested in the so-called STM books or scientific Technical and medical but it can be also cookbooks travel books and so on so books which which have a nature of a reference much more so than immersive Pros such as novels I'm not talking about novels and the specific proposal that I will have is I would like to introduce a notion of tiles which is to some extent inspired by Pinterest um tiles instead of pages this is a very small move but with uh with deep consequences so and I have worked for quite a while with uh with O'Reilly on on their books um in various ways um what was that geez clearly postpone and and that when I was learning uh computer science and programming languages I did it mostly from books but these days people learn a lot from online free resources and in particular I couldn't find it for this conference but but somebody asked a query on quora how should I learn a web design and there were like the top voted answer listed eight different sources but not even one of these sources was an actual book the person was sending readers to online resources such as stack Overflow GitHub and other places but but not books so so this is of Interest also to so so okay um so books are in some sense and battled are by by by free online resources and the Publishers find find it more and more difficult to sell books even though uh this is a curated high quality content right um so what I'm trying to do is to like remove the competition and rather make books collaborate with uh with free resources rather than than compete and lose the battle in the end um so um in in this Century the synergies between books and web content remain untapped so for instance they're for any almost content in O'Reilly technical books there is some discussion on stack Overflow but there is no connection or books which discuss code such as learning spark or Scala have corresponding repositories in GitHub but they are not connected to these books so and code might change independently of of the book book can become stale in the process my first experience with an ebook reader was something I called icarel and this is a large system for in fact O'Reilly book readers written for Android um and and in in that e-reader which has not been particularly successful I have been experimenting with integrating in particular GitHub code with Pages mentioning that code so so that was my kind of pioneering experience which also taught me how extremely difficult it is to do good pagination and this this this topic will will come back but but but pagination really you which is which can be done via tricks with CSS is extremely brittle and I spent an enormous enormous time just trying to to get pagination right with with with varied results um similar efforts similar efforts in the sense of trying to integrate information from the outside of the book into the book are for instance Google Books Kindle to some extent and and there is a company in New York City called cetia which which also has a a bunch of interesting ideas so this was an example from from Kindle I believe which I guess is must be doing some form of named entity recognition because when you highlight like a word in Microsoft it brings two additional Windows which in this case bring back Wikipedia and then a potential translation um however these additional Windows the Wikipedia and translation are like from different domain they they are outside of the book not inside of the book so if you wanted to bookmark something or highlight something brought into into this context you cannot that's another example where a both dictionary and Wikipedia is invoked by Kindle so so this is to say that electronic books do innovate right and and and they do bring some progress compared with paper books because for instance dictionaries can be integrated into into readers but I think this is not enough this is another example from in this case Google Books not Kindle where something similar happens so it knows about Shakespeare and can link it to in this case wordnet and provide some information about the person but again this this this little pain which appears is does not belong to to the universe of the reader this is all hinting towards the solution that I'm getting towards where such information can be integrated into into a scrapbook so to speak of both content from from a book and content coming from the outside since this is Google and Google has Google Maps then then they also have quite an attractive integration of book content and maps and you can go to two maps and here on this on this little pane you also see a link to Wikipedia but when we click it we go to a particular link which opens in in this case uh Google Chrome again moving us away from from the reader so you cannot bookmark it you cannot highlight it it it's it's a dynamic connection but but it's not within your Universe of documents that you're working with So speaking of of pagination I mean pagination is clearly a legacy inherited from paper books and it has it has some some goodies for instance everybody is familiar with uh with with pagination and it can be used for reference something was found on page 54. um and also it gives you a sense like where where are you in the in in a book so these are these are the pluses of course but but the problems with paginational electronic devices are have to do with the fact that screens differ enormously in sizes resolutions you can display text in portrait or landscape modes so the automatic rendering produces sub-optimal results type tables can be cut this this is something that that a human editor would not allow in in when laying out content on paper and also when I'm talking about geometry benefits many people have visual memory and would remember for instance that such and such content was on the left hand side somewhere uh up on the page but but this position really depends very much on like the size of font I mean if you increase the font the number of pages increases everything is repositioned completely so this this is an example where where this you'll get used to it in time is moved completely to a different position when when the text is displayed in a bigger in a bigger font and this is this is an example again of a book which I which I bought I think it could have been Kindle but but again if you are changing fonts then some tables get cut into in right it's not it's not nicely laid out so Battery Nation runs into into problems and uh I have been also inspired by by this guy Peter Myers an Innovative thinker in uh in in the uh area of electronic reading and like everybody has been clamoring for innovation in the e-book space but it has been been all that much Innovation really happening because I mean ebooks are still very much um continuing in the legacy of paper books but this this book published by O'Reilly was was like Peter trying to various various various ideas and he himself is working for citia I think I believe he's still there which is a startup in in New York City which is trying to re-imagine electronic readings so what's what city is doing is similar to what what I'm proposing with tiles but they have a notion of the so-called cards which for whatever reason are square always Square I don't know exactly why but maybe maybe Square when it's rotated doesn't it remains a square um but books have have to be specifically edited for this for this format and I believe that that their cards are still an arbitrary unit of content which is still much more related to to display properties than semantic properties of text so you can you can take a look at city and and see what it what it means So currently when you are when you are reading in Pages this is really linear reading that is if you're on a particular page there is only one next and this is next to the to the right so it's it's one linear but but on a but on an electronic screen you can have something above the page or below the page and and these spaces are conceptually not not utilized right so in addition to having something next you can have I'm thinking maybe related material above maybe users annotations below but but there is definitely a potential of of 2D reading or even probably even more a dimensional reading but but let's not complicate things just now so so the proposal my proposal in topic stream is to introduce the so-called tiles Allah Pinterest and forget Pages all together rather divide divide content into meaningful chunks which correspond to subsections of books with much more self-contained aboutness and and completely forget the notion of of paging with their you know inherent problems of of rendering um and another reason to do that would be to introduce the integration of content from from a book and you can think also about documentation of of some products and whatever with content coming from outside of the book so tiles like I said are about some particular topic or set of set of topics but are not arbitrary as Pages or or these cards in CTR they are much more natural to relate among one another and they may also be easier to read in the sense of like shorter attention spans right people may want to say oh I've read it I know what it is so I was thinking also in in terms of bringing the the French revolutionary ideas here of fraternity liberte egalite or trying to liberate content from the one-dimensional sequence so like I said before in paging there can be only one next but when we liberate um content from one dimensionality there can be several different nexts next in the sense of the spine order of a book but there can be other other next going to complementary material um yes there's another aspect of this of this liberte here namely the fact which is not yet implemented in Google books or Kindle that you can navigate to another book without explicitly closing the current book and opening another book right because what you are really browsing is a graph of tile rather than individual sequences of pages so the liberte here is illustrated by like two different streams so these could be two different books and and there is a possibility of non-linear reading if in particular the fourth tile is somehow related to the first then you have you will have two next you can navigate to the next in the book order or to something related which can be later in in in in in the book the egalita is also extremely important here that is all the tiles are created equal which is even if they come from completely different sources they are treated by the platform identically so all of them are searchable all can be bookmarked all can be highlighted in the same same way so in this in this case we say that that this this the the tiles which are which are surrounded by this by this curve are all treated the same they have some IDs in the system and no matter what their provenance where they are coming from they are treated identically from the point of view of for instance searching so this is an example of of how how the Prototype which is which is an Android application looks like on a tablet and I couldn't figure it out how to how to show you a live demo connecting this tablet to to a computer that's why I took a bunch of screenshots instead um so so this this shows you how a tile looks like and in the left upper corner there is a there is an icon suggesting a stack Overflow in this case and this is a a book which is synthesized from from stack overflow and on the right upper corner there are icons representing books or other resources are the tile sets which are found by the system to be relevant and not just relevant but but but really helpful and there is there's this V thing is an affordance which which shows that there is something about the page then you can kind of slide it down see related material to the tile currently displayed so this is an example of a tile which is filled with with code from from GitHub so I can I I have software which which can create ebooks out of GitHub repositories also eBooks from stack Overflow and kind of an obvious extension to do ebooks containing you know Wikipedia articles and so on and as we go from tile to tile this this right upper corner related material changes dynamically at all times right so for instance here we have three different resources which which are deemed relevant to the tile currency being read and the fraternity aspect is is is really saying that the tiles can form a graph that is that we can have additional edges between between uh tiles treated as nodes of the of the graph and the rest of the topic stream platform is really about Edge discovery that is which which address are really going to be helpful and um this helpful notion is very important because for instance Facebook has introduced some time ago related stories so you may have a story about something and related stories but I find these related stories very often repetitive like yeah the same story told by different newspapers and typically I wouldn't waste time to go there but it would be really nice if if these things were not not just related but in some sense either a contrasted view or some explanation of but but but complementary and and really helpful right so I described this already and and this this fraternity here also tells you that that you can switch from from tiles from from book to another book very easily like I said without explicit opening and closing so um Google Books introduced something called the skim mode which at some point scared me that they have done the same thing and this this is the this scheming mode so if you if you tap on a on a page it can provide something which looks like tiles but what it really is is kind of zoomed out view on which you can see pages to the left and to the right and relatively quickly browse through a succession of pages to find the one that you are more interested in and then you can zoom in onto this page but even though they do look like tiles a little bit they are still pages with with kind of arbitrary assignment of content to visual units that pages are uh so so this is also how this prototype looks like you can have you can have a bunch of books some of these books can be synthesized like I said from other resources so not not the actual publish books but but books quasi books built from GitHub stack Overflow Wikipedia and so on when we go into into a book like this a a pretty standard table of contents is displayed but when you go inside what you see is a tile so not not a paginated system and and again the related tiles are seen on that on the right right upper corner when when we slide down this related material which in this case was uh is is code on the left corner right now because it this this panel moved up moved down you see something like a breadcrumb which says that we are in this book to the left but but we have ventured into into a different tile and and likewise you can you can go and and browse other related material and as we go from tile to tile like I said these whatever is related to is computed dynamically well it's not computer Dynamic it sits in a database but so it was computer Dynamics somewhere on a server download it into the application so the application now knows the address between between tiles and this is also this is an example of a search engine which is which is also running in this Android application locally so in this case I'm entering something called application permission maybe but but with stars representing the rest of the words and this will be the first result because you see whatever what was matched in the in the various books but when you when you click it by a particular book like the first book we also see in which chapters it matched and and you can go into a slide not slide but but a tile where matching words have been found but and all the tiles are indexed independent of of the resource um I'm using a lot of ePub all all around or I don't know if you know it what ePub is but but it's a kind of a popular standard for electronic Publications electronic books which is basically a zip of HTML CSS and and the usual things uh in a in a standardized way okay so I do have a software component which can take a regular book and tile it and then put it back into into it into another ePub but but but now everything has been divided into into small chunks and then in addition I would I can put also uh the content in XML binary format and also components of of the search engine so there's quite a lot of Technology involved here that uh I'm not talking I'm not going to be talking about today but basically that's what I was just saying and this this remains to be a problem that is how to really find which tiles to consider as being helpful and complementary and relevant and and I have started a an implementation using Apache Spark and thank you to marechology here who helped me think through the process of of how to use Sparks rdds to you know match ebook Concepts and tiles into into this process of finding related tiles but of course this process is not not finished yet so I'm hoping uh at this conference even to speak to some of you about the NLP techniques that could be used to actually find discover address I suspect that I will also need a specialized tokenizers for the different domains for instance if we are working with code such as GitHub code I might want to look for classes and see like Java classes color classes and and see whether the same entities are mentioned in text Maybe uh use this 90 to recognition here for relating tiles so this this is not yet really done but but it's a very interesting problem and in particular when when tiles when books are split into tiles there is quite a bit of information available uh the metadata about a tile such as its title and also its ancestry right that is a tile is might be a subsection of a of a section that has a title and which is a part of the chapter which has its own title so there is a whole ancestry path of of of titles and in particular I have a book about Corsica and very often there are sections entitled for instance where to stay where to eat Etc but it is important to know for instance whether this where to stay is in a chapter about bastia or ajaxio right because and so because this this uh parent title really determines like what what is the aboutness underneath and and also the content of the tile will also be a source of features to be used in matching matching tiles but I'm not so much interested in clustering tiles but rather for each individual file finding what should be shown to users to be really most helpful so this I I guess I just described and uh the future work is also I'm thinking about a server uh which which can be an ad server or um a server which which can generate whatever is related and and give us information that can be downloaded into the app about what is related and like I said this was really intended to to work with scientific Technical and medical books but can be used with for instance product documentation and I have actually filed a provisional patent on on some of the ideas here and and I would like to make more people interested in collaborating on this subject so that's pretty much it that I have to say I do have a a demo running on Android both my smartphone and tablet which which I can show you and that's it do you have any questions all right well thank you very much