Devreal

data.bythebay.io: Pauline Ng - Breaking Down Paywalls for Online Health

data.bythebay.io: Pauline Ng - Breaking Down Paywalls for Online Health

Recording: data.bythebay.io: Pauline Ng - Breaking Down Paywalls for Online Health

so 72% of Internet users look for health information online and 26% of these people actually hit a pay wall where when they want additional information they're asked to give more money and if you extrapolate the number of Americans and internet users that means that 54 million Americans are affected um what is the price of a pay wall if you want to read a scientific article it can cost on average $32 to get that scientific article um and the cost of literature has gone up it it has um outpaced inflation by over 250% so it's a huge problem you can get an abstract of of a paper for free but the abstract does not tell you the methods the data or the results and that's really what you need to understand a paper um why would someone need a scientific paper why would someone want to read scientific literature um one thing is well what if you know someone who has cancer or you yourself have cancer or a a serious disease and you want to know details like should I pay more for a special cancer treatment how would it affect my survival if I pay $50,000 more for the specialized cancer treatment another case would be what if you want to find out things that you can't ask your doctor like well making smoking pot make me stupid now that might be legalized soon but you can imagine that there are embarrassing questions that you don't want to ask your doctor but you want to know the side effects for or you want to know what studies have been done um and you can't ask your doctor so if you were to search the the scientific literature to find details about these questions you're going to reach this you know a c something like this where this item requires a subscription to or they're going to ask you to pay a certain amount of money um here are some other examples where scientific Publications are not just for academics I'm an academic I'm an academ I I'm an academic but this athlete found her own genetic flaw based on her symptoms she researched the literature and she was able to find What gene was the cause of um all of the symptoms that she was seeing um this singer was able to diagnose her own ectopic pregnancy um this is Vanessa Coulton and then um this 15-year-old used all of the Open Access literature out there to create a pancreatic cancer test so I'm just going to talk more about the middle case because actually she she was a favorite singer of mine this is Vanessa Carlton um she sang A Thousand Miles I don't know if any of you guys heard of her uh but it's like one of those songs you always sing along to um but she had an ectopic pregnancy she knew it was a topic and she was being treated with methotraxate a drug to save the pregnancy and then she started to experience these weird symptoms um and she was constantly standing up which isn't what you really expect so according to her blog she basically researched herself and read found an article in the lanet which is a medical journal and basically diagnos diagnosed herself to go to the ER and unfortunately her her pregnancy was terminated but she could take that action and I want to point something out um about Vanessa Colton okay so she's you know Grammy Awards winner but she never went to college her education was at the school of American Ballet so uh I don't I don't know like if it's high school and then this but basically there's no college education here and yet she could use this information to diagnose herself so normal people can use scientific literature to really enable themselves and I think it's really important to understand that when something is wrong with you internally you will search and find out and and try to explore and really try to discover what could be wrong with you but this is where the pay wall can really hurt you so what's the Injustice of closed access papers the average research project that the government funds is $450,000 so a scientists do all this research and at the end they want to publish their research how much does it cost to publish research we know how much the publication costs are based on Open Access journals it's about $2,000 so at the final stage of all of this um research that is paid by taxpayers um the Publishers do a little bit of work you know in terms of labor cost and that locks up all of this data now there's um there are actions to make scientific literature um more accessible so the White House has a directive to uh allow um access to publicly funded research but this is a White House initiative and so that's only as long as Obama is um in the White House they're trying to make something permanent but that's in the Senate we don't know when that'll pass there's been several iterations of that so um even though um ferally funded access uh sorry federally funded Publications um are now being open access this does it's not unclear whether this will apply to the Past research 95 years of past research before copyright law came about and a lot of the clinical trials in medicine are paid for by private companies and these are drug comparisons um so they're pretty useful information so privately funded research is still closed and you know this evolves the drugs that you take um the lot of the paper Seekers are not academics so we actually looked at uh Reddit um and we there's a Reddit scholar thread where people can request papers and you can see the topics that they're requesting and 46% of the requests were for Health Fitness and medicine and um in in if you're in a UD University setting people like to publish in high impact factor journals that means that a lot of other academics read it but it turns out what people are requesting are low impact factor journals 62% have impact factor less than um less than five so what are the topics that they're requesting these are just a a sample of paper titles um that you can see from Reddit scholar this is um MDMA so ecstasy and heighten oh heighten heighten stress in recreational ecstasy users something you probably can't have a conversation with um with your doctor doctor stress perspectives on its cognition um and uh Feer fermentation homemade beer um all right so there's a hashtag I can has PDF where you can request a paper and um someone will get that paper for you and email it to you and this has had extremely viral growth I don't know if the colors are showing but it started in um January 2011th and you can just see that it's this is actually a break so it it it's um it's an exponential growth here so a lot of people are requesting scientific papers and need scientific papers uh I can has PDF was actually addressed by scub this is a website um run in Russia where they um where a woman is accessing the true papers and putting it uh in a in a repository now this is this violates copyright so it's dark Hub they have a lawsuit um and there's another similar website called library library Genesis but this is like the second generation there's um over 40 million papers here and these are full papers so while a great resource it is violating copyright um the advantage is you can get the full text you have the full paper the disadvantage is we don't know how long the content will be available links can be taken down and um they can be temporary URLs and it's dark web so it you you have to access it through the scub website you can't NE Google won't you know if you search for those topics Google won't won't lead you to that um and popularity attracts lawsuits so the more popular it has become the more likely it's attracted lawsuit and Elsa has already sued um scub um this example is actually uh author are putting up their own papers their own papers and um they're being sued to take it down so um authors themselves can't um put up their own papers so I want to talk about a solution that we're proposing called fact Pub and aderes to the principles of copyright uh first we have to talk about what is copyright so facts are not copyrightable um and this was established in a in a lawsuit where um the questions can you copy white pages but white pages are facts and so um this CA this uh case uh they decided that facts are not copyrightable and if you think about it a lot of scientific Publications are just composed of fact um recipes are also excluded from copyright protection um and this was a case of a Dan and yogurt recipe book um and if you think about it meth methods is very similar to recipes right recipe is you do this you crack an egg there's five eggs you crack an egg you put it in the oven for 350 degrees methods in the scientific paper are you use these components and you do it in this order um equation equations figures and texts are the uh only ways to express and so these expressions are not copyrightable so figures is a little bit tricky but equations are facts because they describe something inherent like equals MC squ that is the behavior of physics and so that is not copyrightable text is definitely copyrightable though so mentioned here can immediately put in question the whole statement so so in this case it was it was the only way to express this idea like in this particular lawsuit the um the com the plaintiff who was trying to sue for copyright could not come up with alternative ways of text to explain their idea so if it's the only way to express an idea then then it's not it's not copyrightable I I would say the equation is is pretty solid but the others are definitely more expressive um and all of the cases that I've cited fall under this idea of uh idea expression divide which is this doctrine that ideas are not protected by copyright but the expression of the idea is protected by copyright so going back to the recipe example five eggs break the eggs you know crack the eggs that is not copyrightable but when I picked the five eggs from my mother's farm and I felt the warmth of the egg you know in my hand that is an expression of of five eggs and so that would be copyrightable so what's the idea all right scientific papers uh the readers want to access the scientific paper but it's blocked by um a pay wall academics now actually uh have access beyond the pay wall and they can download papers on a local computer what we provide is a uh a Java Standalone that locally extracts facts from the papers and then this this standalone will distribute the facts to our web server so what happens is this a repository of scientific facts where um the academics can access it and the public can access it um and one thing is that the academics because they have uh if they log in and make an account they can also edit the facts okay so the key here is what we're doing is that a paper donor sorry a fact donor distributes facts and not not the paper this is just an example of our um Java it's a standalone and you would install it locally on your own computer you just drag and drop a couple of papers and what's happening in the background is that facts are being extracted um and the source code is available at this GitHub and we can also run it by command line okay it takes about 30 seconds to process per paper um and it depends on the number of pages on the paper oh I wanted to show you a demo so um let me just go to the demo right here and I'll get it running so this is the fhub website where we are um hosting everything and this is if you are academic and you have access to papers you don't have to be an academic you just have to have access to papers so idea is that you would download this tool and I've already downloaded it um just to save some time okay so now I'm going to execute it so um oops it's already open let me just okay so there it is and um I'm just going to drag a paper here I'll show you the paper this is the paper um it's called it's a random paper I don't even know what it is a human population responses to talk toxic compounds and what's going on right now is it's extracting facts and it's going to upload it to the server so um while it runs um I'm just going to continue my talk and then we'll come back to it in a second or so okay so so it's extracting facts um and then after it extracts facts you can click on the link and it'll take you to the page uh where the facts are extracted oh here's the demo it's running right now okay so how does um the fact extraction work so um the first thing is that we take PDF and we convert it to text uh we extract texts and tables and this is actually using the work of clamp at all and it's actually kind of challenging because of the structure of a paper it isn't like pure text you actually have to recognize when there's tables and when there's figures and that's why we use existing software to do that and this is our cont contribution where we now take the text and we use natural language processing to convert the text to facts and we made it really light and it's basically 10 um Simple Rules so that it can be domain independent and we also extract acronyms so we have a Java Standalone and we decided to release that because the required libraries and tools were already implemented at Java especially this one and Java is crossplatform so people with Macs and um PCS can use this so this is just an example of transforming a sentence into a fact so in this paper um it mentions Two drugs treated animals um uh perform significantly less well on the road than the MC treated animals and then what you can see the fact here is a condensation of the sentence here um where the uh two drugs treated animals significantly Less on the road or than MC treated animals so you don't get everything um it's not it's not perfect you're going to have to infer something it's um because we can't copy the sentence directly we're trying to get the uh important aspects of the sentence yeah this is horrible it looks like instead of animals performed less it looks like they were treated less the meaning changed see it says treated animals performed significantly less well than down below it says treated animals significantly less that's just different we that was our first version where we have aimed to identify important verbs um the thing is we figure that if you really want the paper you actually have enough information to know whether to buy the paper so every single sentence is converted into a condensed version of a fact and if you look at it enough and if you're in the field you can kind of get a sense of is this what you want to read then you could go and buy the paper um but it's there at least so that you don't pay $30 and then read the paper and decide it's not what you want we actually have um improved that um that was a version uh that was like version one and we've now added verbs but we haven't reassessed um the performance so these are 10 rules for fact extraction how much time do I have left doing 30 minut 30 minutes I have okay I'll just I'll just keep going okay so um for the first version of the algorithm we had nine users examine 556 sentences and their fact representations and 70% of the facts were rated meaningful the rest were um rated misleading or meaningless and that's that's an issue right so it everything you can't trust everything on the website but the data is there in a condensed fashion for you to kind of see what that paper is about um results um if we broke it down by section results highest in terms of the percentage of facts rated as meaningful we missed a lot of verbs because of the way our we implemented our rules um and that was as low as 64% we our second version we've um Incorporated more important verbs we'll have to see if this improves so like I said before we were missing important verbs for example patients in group a died we kind of missed died and that's very important and we've we hope we have identified that in the next version um essential adjectives for missing um gastrointestinal tumors for example and important prepositions so we're working on these two and um our next version we'll have these twoo okay so the other challenges is um you know this is an automated process and we've tested on a few thousand papers but um the parser from PDF to text can fail we also take a lot of meta information but we have to infer the title or DOI from uh the PDF and that can um be missing or it it's just erroneous or go wrong okay so I've talked about uh the fact extraction and now I'm going to talk a little bit about the public website hopefully this is done now okay so this is done um processing and so now you can just click on here um and you can see uh what was automatically extracted was the title and um like I said we try to identify the title or the uh it's a DOI is a digital object identifier which tells us a lot of information about the paper and um once we have it we can populate this field with all of the author's title and um um Journal information and you can see this is basically everything this is the abstract of the paper which we can copy directly these are acronyms that we've infer uh oops here acronyms that we've inferred from the um from the text um looking at capital letters and then every line here represents a paragraph and every bullet point represents a tech a sentence that has been converted into a fact um and if you read it you can kind of get a sense of what the paper is about okay so I just kind of I just kind of describe the fact pub website um the engine behind the fact pub website is media Wiki um and we've processed 3,000 papers and the idea is that people would download the Standalone um and run it on their own computers and it scales with fact donors um and that way the fact donors don't violate the law because they're not Distributing papers they're Distributing facts oops and I already returned to demo so um to prevent spammers because we had a lot of problems with spammers is that you if you want to edit the paper because you know there are errors like you noticed some verbs were missing and some were meaningless you actually have to create a a a an account and then um when you log in at the Java on the Java Standalone It'll recognize it and then you will have permissions to edit this paper um so only the paper can donors and the people who possess the papers can edit the facts the other thing is we have a search box so if you're interested in cancer or any topic then you can search it obviously we only have a few thousand papers now and with um you know over 50 million papers out there there it's very small but we hope it will grow um as I mentioned before we have to infer the title or DOI in order to get paper details and then we um we have to uh we also extract acronyms and tables and then um facts from papers so um each sentence is converted to a fact the each paragraph is delimited by lines and then we also extract the section headings for organization so all of this will not substitute for the full paper like you say it's much easier to read the full paper it's sentences it's what we're used to but there is enough information to decide whether to purchase the full paper and we hope there's enough information to make a decision um whether to purchase the full paper or not at least you don't have to purchase it and then decide this isn't you know this wasn't worth it or this doesn't have the data that I that I don't that I wanted so what's next um we'd like to um transform this information it's really a subsequence of the sentence and so the more we can add to it and modify it um the more we create new content so we can modify the text and rearrange the text what we really need to do is engage users so right now it's a Java Standalone and and we figure people maybe they'll download the Standalone and they'll you know upload papers once but how to keep them engaged in upload papers monthly or or um or actively because it really scales with users um maybe we can mail updates or list top contributors maybe gamify it um maybe we can um you know honor the paper requests that tweet out to I Can Has PDF um the other thing that we've thought about which is which is actually a huge commitment is to convert the Java Standalone to JavaScript because there's a lot more um natural language processing tools um and all of the tools were in Java it would require a lot lot of work to go to JavaScript but the idea is then people could um install this as a browser extension so that whenever anyone browses a scientific paper it would automatically send it to the fact Pub server okay so fact Pub is a permanent repository to provide scientific facts to the public so that the public can search facts and find them and the idea is that if we have enough information and it becomes Big Data where there's enough um facts in there then it could Aid in training um decision-making for automated Health decisions um but it's really relies on crowdsourcing to populate the database um so here it is um and you can go to the URL F hub.org uh also I have a position available in Singapore we don't pay much we don't pay as much as Singapore uh sorry as San Francisco but it's only like $150 to go to Thailand or IND you know Bali so there's there's that Advantage there that's it [Applause] thanks yes um so this is really interesting and I've always been as someone who tried to get pay yeah ridiculously expensive um I'm wondering you know it seems like this is a similar industry where we're in this base transition right where there there was a reason existed they took on certain amount of effort paying their bills but as as publishing has become easier and Tool chain there's every day it seems like there's new ways to publish things faster and easier especially in the scientif community and um so is there is there like a root cause like is it is this really like at the end of the day this should be something that is ultim i s the mention the White House's efforts but is this really a play that this information is in it's for the public good and that we should we as taxpayers will just bundle that with with the research right we'll just it's so cheap to publish things we'll just do the little administrative stuff to get that stuff out there for you and then that and then that you don't have to worry about think different I I don't know is there another is there sort of yeah a lot of universities systemic changes that need to happen I don't know if it's systemic but like some universities like MIT require that when you publish a paper that you put it in the MIT repository that it's open so even if you go into a closed access Journal it will eventually become open um but it's not systemic like the White House initiative that exists and that's a temporary thing but the law um even then you know the White House was six months after publication it would be free to the public but already they're negotiating for the law that hasn't been passed yet that'll be one year so you know there's a lot of kind of we'll see what happens um the thing is the Publishers have this huge Treasure Trove of data that they don't necessarily want to give out to the you know like that it's a lot of money right and if you need that data you're you're willing to pay for it or the every University is paying tens of thousands of dollars to get access to those journals so so they're pretty powerful lobbers I mean just like you have Universal uh Paramount Pictures or the movie you know if you look at people who are very much in extending copyright you've got Disney you've got the movie um industry and then you see you see a couple of AC you know scientific Publishers you you see a couple of those on the list is there while it's not ideal uh to repay is there is there a role that for libraries can play here to make that information more accessible um I mean libr already do that to a certain extent is there a way to think that easier or better or more I guess the the idea is that you could go into a University Library if they don't have security measures and then maybe look it up but you know some places you required a school ID to go in the other thing I've noticed is that um some of them have responded by lowering the price so they say $5 for Access like you read it I don't I've never paid it but it's like $5 and maybe you read it but you never download it and it's kind of like I think it's their way to make the price lower so that there's less complaints well uh I would surmise that uh calling this facts is mostly a way to circumvent the copyright protection because in the scientific papers there are a lot of statements that could be mistaken for facts yet they're not a facts like a hypothesis that is stated and then discarded or disproved and rejected and yet that hypothesis when taken out of context may appear to be a fact even though it's really a false also an aome or a postulate or a theor posed as an open problem is not a fact in fact the prove of a theor the pro proof of a theorem is hardly a m fact because contr to stating just a known fact a lot of actual work creative work goes into proving such a theorem so misrepresenting that as a fact I would uh on my part the objective we have a disclaimer so on the website we do I mean we know we already know that 30% are wrong so on our website we have a disclaimer um and we can certainly make it more you know obvious that do not take this as I mean it you know it's it's a mess of words where you get a flavor of the paper I mean when you start reading it if you're in the field you kind of get it you know if you're in the field and you read the paper or if it's kind of tangentially related so this this tells you that um the information can be inaccurate and stuff like that we're also working to identify opinionated words um um and also like negative words ju to um to see to catch those phrases and we might just throw out those sentences completely so this are basically statements statements extracted from text right the facts are basically not necessarily true facts right statements logical statements exed through text which may not meaning I mean it's a very yeah understand pretend that these are facts but some of these are facts and some are not facts so it's sort of like OJ Simpson showing up and telling them give me my stuff because they were going to auction off some of the property that was previously stolen from him and yet some of the things that they actually grab where somebody else's baseball and such and now he's serving his term in prison because he committed uh what a bing right I have a question uh so uh what were the main kind of customers customers yes users users so we first need contributors to be honest right we need to get to a certain we need enough data and um yeah so it it'd be what we're seeking first is people to use it and get feedback and how to make it so that people continually use it and not it's not just a one-off and then that would be the first phase and then if we get enough in there then ideally the public would be using it when they search for a particular paper and say they can't find it they would hit this and then they would skim and decide whether they need to purchase the full paper but then even beyond that if we got enough papers and if we've got enough facts um I'm I'm actually in the health field of computational biology if you think about um like decision-making for health decisions and if you want to automate it like if you think about um go Google's go they actually had to train on a database of go uh games um and go positions and stuff like that so if you had enough papers and if you had enough facts yes some of them could be wrong but the idea is that in if you had enough then hopefully the noise would go down or you'd verify it with you know with like um person assistance um AI um but then you could probably build that feed that database into um like one like maybe IBM Watson or something like that to see if any of those facts help with diagnosis or something like that cool so how can I mean you know a lot of this conf is about community so we really want to surface things like this to the general public right because we have a life sciences day on Friday where multiple data sour medical test so uh I think this is multiple kind of you know this Stu have been on Friday as well right so we can multiple purposes so how can we engage uh authors to submit more yeah I so we need ideas for that and you know I have um I have like a couple of programmers doing it but um these are some ideas that we have I think we need like users and telling us what's wrong and how do we engage and priorize our things because like to do this to make a JavaScript browser extension which was the initial idea it just take an immense amount of time and effort so the question is can we just engage users with the Java Standalone and I mean like I wish I could gamify it but I just don't have that mentality like and or know how to do it um the the Java source code is available on um GitHub uh the link is in one of the slides there and just your comments um contact me and um I I mentioned at the end that we have like very very low paying internships if anyone wants to you know use Singapore is a base to explore Asia thanks