Devreal

Keynote Address: Now is the Golden Age o...

Event: Text by the Bay

Text By the Bay 2015: Mark Liberman, Keynote Address: Now is the Golden Age of Text Analysis

Recording: Text By the Bay 2015: Mark Liberman, Keynote Address: Now is the Golden Age of Text Analysis

okay now is the golden age of text analysis or will success spoil computational linguistics so as we've heard from the previous two speakers a new evolutionary niche is opening up for text analytics why well it's pretty simple there's a digital shadow universe which increasingly mirrors real life in the flows and stores of bits and society is mostly about communication and communication is mostly text or talk which is just text and fancy calligraphy and more and more often this exists and these this flows around and gets stored in digital form they're simple properties of text like the words that make it up that are a pretty good proxy for content um what we can do with them is far from perfect but it's not bad and this is at least better than anything else we have at the moment and bigger faster cheaper digital everything and better programming languages and some improved algorithms and so on make it easier and easier to pull content out of the flows of text in that digital shadow universe so there's an old argument about whether content is king or communication is king but the content of communication is at least the power behind the throne so in that new evolutionary niche there's a host of newly evolving life forms that have got the means and the motive and the opportunity to live off of these flows and stores of text and of course they add their own digestion products to the ecosystem so these text rats are flourishing and the ground sloths and aurochs and sabretooth tigers and other emerging digital megafauna are beginning to grow up but you know all of this the cenozoic metaphors aside or you wouldn't be here but there's something that many of you may don't know or may not know something that's played a critical role in creating the world that you're moving into and i'll explain it in the form of a story and this part of my talk is based on a presentation that i gave a few months ago to a workshop called statistical challenges in assessing and fostering the reproducibility of scientific results which is a mouthful run by the committee on applied and theoretical statistics of the board of mathematical sciences and their applications of the national academy of sciences and that workshop was pretty alarming actually there's really a crisis of credibility in many areas of scientific research as documented elsewhere before and since for example i don't if if any of you haven't read john euanadise's 2005 paper why most published research findings are false published in plos medicine i recommend that you look it up on google scholar and read it the chronicle of higher education a couple of months ago published an article under the title amid a sea of false findings the nih tries reform and here's a quote from that article als researchers seeking a cure for lou gehrig's disease went back and reproduced studies on more than 70 promising drugs they found no real effects zero of these were replicable dr francis collins said he's the head of nih zero and a couple of them had already moved into human clinical trials what i'm going to tell you the story i'm going to tell you just concerns a different crisis of credibility that afflicted a different research area 50 years ago but along the way as a result it created human language technology as we know it today so once upon a time there was a bell labs executive named john pierce he supervised the team that built the first transistor and he oversaw the development of the first communications satellite so credibility was not a problem for him that's john pierce in 1966 he chaired the automatic language processing advisory committee known colloquially as alpac which produced a report to the national academy of sciences under the title languages and machines computation in computers in translation and linguistics and in 1969 john pierce wrote a letter to the journal of the acoustical society of america that was published under the title wither speech recognition and neither of these interventions were friendly well we'll see so the alpaca report noted that machine translation in 1966 was not actually very good although it was the result at that point of 10 years of quite large investment by the us government which thought that you know the russians are out there competing with us and we don't have enough people who can read russian technical literature and so we'd better get computers to translate russian into english so that our scientists can follow what the russkies are doing and so you know russian english was sort of the main thing so the alpaca report said that the committee cannot judge with the total annual expenditure for research and development towards improving translation should be however it should be spent hard-headedly towards important realistic and relatively short-range goals which was a sort of polite committee-speak way of saying you know this is don't support it the committee felt that science should precede engineering in such cases they hadn't read peter drucker who actually hadn't probably written his paper by then we see that the computer has opened up to linguists a host of challenges partial insights and potentialities we believe that these can be aptly compared with the challenges problems and insights of particle physics certainly language is second to no phenomenon and importance and the tools of computational linguistics are considerably less costly than the multi-billion dollar accelerators apart part multi-billion volt sorry also dollars these days accelerators of particle physics the new linguistics presents an attractive as well as an extremely important challenge the funders read between the lines and u.s government funding for machine translation went to zero for more than 20 years and it wasn't transferred into linguistics i can tell you john pierce's views about automatic speech recognition were similar to his opinions about machine translation and his 1969 letter to jasa was much much less diplomatic than that committee report because he could just write it and get it published over his own name with her speech recognition a general phonetic typewriter he wrote is simply impossible unless the typewriter has an intelligence and a knowledge of language comparable to those of a native speaker of english most recognizers and by this he doesn't mean speech recognizers he means researchers working on speech recognition most recognizers behave not like scientists but like mad inventors or untrustworthy engineers the typical recognizer gets it into his head that he can solve the problem the basis for this is either individual inspiration the mad inventor source of knowledge or acceptance of untested rules schemes or information the untrustworthy engineer approach the typical recognizer builds or programs an elaborate system that either does very little or flops in an obscure way a lot of money and time are spent no simple clear sure knowledge is gained the work has been an experience not an experiment and then he went on to tell us what he really thought we are safe in asserting that speech recognition is attractive to money the attraction is perhaps similar to the attraction of schemes for turning water into gasoline extracting gold from the sea curing cancer or going to the moon one doesn't attract thoughtlessly given dollars by means of schemes for cutting the cost of soap by 10 i'm not sure that this is actually true as a matter of investment philosophy i think if you could reliably guarantee to cut the cost of soap by 10 you could get funded but anyway to sell suckers one uses deceit and offers glamour it is clear that glamour and any deceit in the field of speech recognition blind the takers of funds as much as they blind the givers of funds thus we may pity workers whom we cannot respect so that's pretty strong stuff the fallout from these blasts well the first idea rog ready and some other people at cmu this is back in the days of classical ai when people thought that artificial intelligence was applied logic they said well all right yes all right so it's true that in order to do speech recognition right you need intelligence you need to understand but we know how to do intelligence and we know how to understand content so let's you know give us a shot and so darpa organized something called the speech understanding research project 1972-1975 which used classical idea classical ai to try to understand what is being said with something of the facility of a native speaker as pierce has put it and this project was viewed as a failure and funding was abruptly cut off after three years basically because the internal management at darpa had had read and assimilated john pierce's letter to jasa and when they did the first demos of the was the preliminary results of the project they were not sufficiently impressed to continue funding it and the second idea was just give up so between 1975 and 1986 there was no u.s research funding for machine for either machine translation or automatic speech recognition now john pierce was not the only person in the world much less in the united states with a jaundiced view of r d investment in the area of human language technology by the mid 1980s many mainly maybe most informed american research managers were equally skeptical about the prospects for progress in that area but at the same time there were many people who thought that hlt was really needed and maybe in principle was feasible so they had kind of one and a half of those three circles in the venn diagram that we saw a few minutes ago but all three circles were definitely not there so in 1985 the question arose should darpa the defense advanced research projects agency or arpa as it's also sometimes known the advanced research projects agency without the d restart human language technology and charles wayne who was on loan to darpa from the nsa had an idea to design a speech recognition research program that protects against glamour and deceit because there are there's a well-defined objective evaluation metric you're not going to do demo you're not going to do evaluation by demo you're not just going to send some generals and admirals around to see whether they can be whether you can whether the researchers can impress them with their public relation skills rather the national institute of standards and technologies will actually apply a test that's been defined in advance on shared data sets so you're this is not a benchmark on some unknown program so to speak it's a benchmark on something that's well defined and you'll ensure that simple clear sure knowledge is gained because the participants are required to reveal their methods to the sponsor and to one another at the time that the evaluation results are presented at a workshop of some kind they people weren't required to share their programs but they were required to be pretty specific about what they did and if other people couldn't replicate what they did they heard about it pretty quickly now this required published data and well-defined metrics so dave pallett who ran the group at nist that that handled the validation for the first couple of decades wrote in 1985 and sort of the run-up to this project definitive tests to fully characterize automatic speech recognizer or system performance cannot be specified but it's possible to design and conduct performance assessment tests that make use of widely available speech databases use test procedures similar to those used by others and are well documented these tests provide valuable benchmark data and informative though limited predictive power there's a very interesting story about the texas instruments digits database and george doddington's role in publicizing both the database and the results of applying it which i will not tell you about right now but you can ask me about it during the coffee break it's a good story so the result was what has come to be known as the common task structure it starts with a detailed task definition and evaluation plan which is developed in consultation with researchers and is published as the first step in the project before anything else happens there's automatic evaluation software which is written and maintained by a neutral third party and published also at the start of the project crucially their shared data training and dev test development test data is also published at the start the evaluation test data is withheld for periodic public evaluations now not everybody liked this idea a lot of piercings including john pierce were skeptical basically their idea was you can't turn water into gasoline no matter what you measure and researchers in general were pretty disgruntled i remember one researcher who later became one of the strongest champions of this kind of research saying it's like being in first grade again you're told exactly what to do and then you're tested over and over they thought they were being treated like children but it worked why did it work well one thing was obvious it allowed funding to start because the projects were seen as glamour and deceit proof it allowed funding to continue because the funders could measure progress over time in an objective way even when real products were decades in the future so you're not the research is not yielding something you can sell or something the army can use right then or even the next year or even after five years but it is producing steady progress on an evaluation metric that arguably is leading you towards a threshold you can define the point at which it could be usable a less obvious reason that it worked is that it allowed project internal hill climbing because the evaluation metrics were automatic and the evaluation code was public for some it's hard to believe in retrospect but this obvious way of working was a new idea to many people and the same researchers who had objected to being tested twice a year began testing themselves every hour or basically as fast as they could as they could churn out results they would diddle some something in the algorithm and try again and they were sort of hill climbing on the problem a third and even less obvious reason was that it created a culture because researchers shared methods and results on shared data with a common metric and participation in this culture became so valuable that many research groups joined without funding um darpa real or the defense department in general realized what they had hold of in 1991 when they held a conference on it was basically the first trek conference so it wasn't called that then and they had funded four groups to do research in that area and 40 other research groups from around the world participated and showed off results without being paid anything for their participation and actually i think the three top scoring projects were among those who hadn't been funded um and so they realized that that having these benchmarks and shared data and open workshops was an extremely attractive structure for for fostering research in areas of interest to them the other thing that it did another thing that it did is that the common task method by its definition created a positive feedback loop because when everybody's program has to interpret the same evidence which is same ambiguous evidence ambiguity resolution becomes a sort of gambling game that rewards the use of statistical methods and has led to the flowering of so-called machine learning um given the nature of speech and language statistical methods need the largest possible training sets which reinforces the value of of open data of shared data and these iterated train and test cycles on this gambling game are i think literally addictive i mean for the same reason that one-armed bandits are addictive people who work in this area they you know their dopamine levels rise in consequence i think of scoring a little bit higher on the internal evaluation metric and this creates the simple clear sure knowledge that john pierce wanted and motivates further participation in the culture so the common task method has become the standard widely accepted research paradigm in experimental computational science not just in natural language processing or speech not just in human language technology and it always involves published training and testing data well-defined evaluation metrics various kinds of techniques to in to try to avoid overfitting managerial as well as statistical and the the domain is basically anything that involves algorithmic analysis of the natural world over the last three decades this method has been applied or variants of the method have been has been applied to many many other problems machine translation speaker identification language identification parsing sense disambiguation information retrieval information extraction summarization question answering optical character recognition sentiment analysis image analysis video analysis autonomous vehicle navigation and so forth the general experience is that error rates decline by a fixed percentage every year to an asymptote that depends on the task and on the data quality progress usually comes in many small improvements it's not the case that there's a brilliant a single brilliant invention that takes you from a sixty percent error rate to a two percent error rate rather there are several hundred small improvements each of which gets you a little bit of the way and actually an improvement of one percent uh absolute um can be reason to break out the champagne and you go to conferences about in many of these areas and in a way it's kind of it's sort of depressing because you hear paper after paper after paper which people did some kind of heroic mathematical and computational and implementational labor and they tested it and they improved the score on a standard benchmark by half a percent um and the paper was accepted and everybody's actually everybody's happy because you know you do that 50 times and you've actually improved things a lot and of course in many areas of human technology this is kind of how things work the things that separate today's automobile from the automobile of 1910 is not one single invention but many many thousands of innovations some of them are more obviously some are more important than others but mostly it's just a little bit a little bit a little bit diffusely moves you along shared data does play a crucial role and it's reused in unexpected ways and glamour and deceit have mostly been avoided human nature being as it is there's been a certain amount of glamour and a little bit of deceit but not very much of either one there are dozens of current examples and some of them are shared task workshops here are a few of them in the last couple years and many many others i i seems like email comes in every day with another half dozen of them but they're certainly dozens and dozens some are just shared data sets and evaluation metrics so text analysis conference is an example they have tracks their 2014 tracks were knowledge based population and biomedical summarization the mcconnell has shared tasks semantic role labeling joint dependency parsing and semantic role labeling co-reference grammatical error correction trekvid which is not really natural language processing but video processing um their their 2015 tracks include included semantic indexing interactive surveillance event detection instance search multimedia event detection localization video hyperlinking etc and each of these each of these is one of these shared tasks these common task structures with a published evaluation metric that's automatic with shared training in dev test data and withheld eval test data another recent example that as far as i know hasn't really had a conference associated with it but has been deservedly very influential is the google street view house numbers data set a real world image data set for developing machine learning and object recognition so there's a whole bunch of stuff 20 73 257 digits for training 26 000. for testing and another 530 000 odd samples to use for extra training data and this is kind of what it looks like that is each of those little squares is a fragment of a google street view image that shows a digit or several digits and the task is to have to write a program that figures out what those digits are here is the performance over four years they went basically from 37 36.7 error percent error to 1.92 percent error and not every uh not every problem shows progress that that's this rapid but this but steady progress almost always happens when you set things up this way and many of you may have seen that this is an old uh figure from showing um pro nist stt speech to text benchmark test history up to about 10 years ago starting from 1985. there's progress is continuing in speech to text so the switchboard corpus of conversational telephone speech which was collected with funding from darpa at texas instruments in 1991 and has been used in literally tens of thousands of publications since then in terms of speech recognition error rate it stalled at about 20 to 30 percent word error rate about 15 years ago but recently it's come down and and it was at nearly 50 word error rate in 1995 uh but recently it's come down to a little over 10 word error rate which is still worse than people but it's sort of getting into the range human disagreement about transcription of this kind of stuff with careful uh definition of the transcription task is maybe three to five percent so there's still a little headroom but progress is being made as of 1983 that is before this process started the first conference on applied natural language processing had 34 presentations none of which used a published data set none used a formal evaluation metric more recent sample in 2010 the 48th annual meeting of the association for computational linguistics there were 274 presentations and every single one of them used published data and published evaluation metrics except for a few that describe new data sets or new evaluation metrics so will success spoil computational linguistics what would that even mean why did i ask this question well today's best human language tech technology algorithms though enormously improved over where they were 50 years ago and quite usable for many practical purposes in fact i think the immediate future is really exciting for the things that we can do the fact is they're still pretty crude we're a long ways from having hal there's a lot to uh a lot of problems to solve and some of them are not easy problems now luckily for us crude methods can be very effective but they're probably really hard problems that are left and i think that the common task method which has been the recipe for i think continues to be the recipe for progress in this area is in danger or is potentially in danger maybe not so much in this crowd but there are there are other crowds first funders no longer really fear glamour and deceit because after all human language technology works so why not just give people money to solve the problems you want them to solve and not worry yourself about all this complicated stuff about task definition and evaluation metrics and shared data and all of that just tell them the kind of thing you want to do and give them the money and and this is we see this happening even at darpa in some cases uh and second obviously commerce often trumps community at some companies researchers won't even tell outsiders what they're working on lyle was telling me about a former student of ours now working at a local company not to be named who explained that he couldn't tell lyle actually what even what general problem area he was working on because it was viewed as company secret and lyle's only opportunity for sort of guessing what it might be was by trying to infer what he didn't say what the student didn't the former student didn't say on the basis of the questions that he asked to give some indication of what he might be interested in so the possible results might be if this if if this engine of progress stalls or you know is abandoned there might be little progress on remaining hard problems or progress with on one of the remaining hard problems might slow down because now you're going to have people in their individual corporate or startup silos working by themselves without the community without the amplifying effects of the of community exchange and interaction and there there might even be a return to some glamour and deceit i hope not the common task culture is pretty strong it obviously has connections into the open access and open source movements and i think that also helps and many companies or at least the people at many companies understand that a rising tide lifts all boats and that there's uh that development of improved technologies in these area in this area will enormously expand the markets in which they can compete with other people this is a picture of george washington's chair at the constitutional convention sorry that the top of that legend went away that's what the top of it looked like and james madison in his diary of the events of the constitutional convention wrote on september 17 1787 whilst the last members were signing the final document dr franklin looked toward the president's chair at the back of which a rising sun happened to be painted looking toward the president's chair at the back of which a rising sun happened to be painted observed to a few members near him that painters found it difficult to distinguish in their art arising from a setting sun i have said he often and often in the course of the session and the vicissitudes of my hopes and fears as to its issue looked at that behind the president without being able to tell whether it was rising or setting but now at length i have the happiness to know that it is a rising and not a setting sun what about text analysis and human language technology in general well i'm cautiously optimistic thank you so uh i was instructed to absolutely finish by 10 o'clock and i kind of hurried through so we have a few minutes our questions allowed alexi okay well in that case questions objections suggestions do we know which way the chair was facing was the sun rising in the very good question to which the ant the answer can be determined because so the question is this gen dimitri uh observed that in fact franklin should have been able to figure out whether the sun was rising or setting by knowing which direction the chair was facing right was it in to the east or to the west of course it could have been to the north or to the south which would have made things more ambiguous i think it is possible to determine that because the room in which the signing took place still exists in philadelphia and i believe that the layout of the room is still as it was then and so next time yes we can google it and find out so please let us know yes well you know i think that there are lots of there's lots of evidence that this method has sunk fairly deep roots in not only the human language technology community but in the visual information processing community and some others and so i i mean one of the things that the kaggle process lacks is some of the community formation you know sort of people getting together to talk about their problems and report their different results and so on um certainly over the years at the workshops run by darpa and by others by connell and so forth at which these tasks have been each year's version of the task have been discussed some of the most important stuff happens in the corridors and at coffee breaks and so on and i think my understanding of the way that kaggle works is that you don't certainly don't get as much of that i suppose that there must be some there are opportunities for communication but they're more diffuse so but but still i think there you know i think there's there's quite a lot of evidence that there are that even if the u.s government were to completely abandon this method which i think is unlikely that there are lots of scholarly and scientific societies and companies like kaggle and so forth that are are sort of keeping it going that's one reason for being optimistic yeah well so um the darpa process has never included publishing code although people sometimes do but it's not required but it they people are required to give a specific enough account of what they did that somebody else could reproduce it and so when someone reports a result that makes a significant impact other people many other people typically immediately try it out and if it works for them then that becomes common knowledge and if it doesn't work for them either because the original success was a fluke or because some things were left out of the description or whatever that also becomes general knowledge and and so you know what you you get a you know this is not a mysterious process it's quite a lot like i don't know what happens in the hot rod culture or something like that that is you have a whole bunch of people who are trying to do things say to improve the performance of souped-up cars and to some extent they want to keep secret sauce to themselves but in fact they're inclined to communicate their methods to others you know on purpose or by accident and so the the information spreads uh well i guess that's true uh that that is that's certainly a question that is uh did this progress happen not because of the method that was used but just because computers were getting bigger and faster and cheaper and one argument against that is that there are areas where for example in the area of image analysis these techniques weren't used for about 10 or 15 years and then they began being used and i think since they began being used there's been much faster progress similarly in machine translation um the the mt people got into the game uh six or seven years in the early 1990s after the speech people had gotten into the game and they were actually they had to be sort of dragged in kicking and screaming by darpa who said do it or we'll cut your funding we won't give you any money and like and similarly machine translation improved rapidly once that began now they're there it's still the case i suppose from a sort of history of technology point of view that there are several different uh theories you could have for example i mentioned the fact that this structure incur encourages the use of statistics and machine learning techniques because you see the process as a kind of you know gambling under making decisions under uncertainty and it's possible that that's that's certainly part of the key and it might be that if people in machine translation had been persuaded of that without having to be pushed into this culture that they would have made progress as rapidly i i don't think so myself i i think that i think the the community so so i think the community formation part is actually critical and part of the community formation is having the same task um so but that's just my opinion i have to say yeah significant the people at cmu have gotten darpa funding since the mid 1980s every year the people at cmu have gotten darpa funding since the 1980s every year some of the people at stanford have participated in darpa projects i don't believe that there's a single one of the stanford nlp program stack that doesn't use techniques that were first pioneered in darpa funded projects oh well in the last 10 i mean the government has not funded any uh parsing research for example in the last 10 years to speak of they've assumed that parsing will be used in other things they have funded and that's sort of indirectly funded such things but sure the the progress in parsing is is now mostly being driven either by independent academics who are getting money here and there or by people at places like google um you know who are doing it for their own reasons and who let other people know some of what they've what they've learned and discovered so i i didn't say that the uh my concern is that government funding wouldn't continue although there's a difference between funding stuff that you can use now and funding stuff that is a decade away or 15 you know some unknown amount of time away from being useful in projects and i think it's hard to imagine the the only peop the only historically the only outfits that you can rely on to fund things with a 10 or 15 year time horizon are the government and monopolies and so for the moment you know google is doing a certain amount of that because they're close enough to being a monopoly um but and they're also corporate culture issues so the you know google is very very very different from apple in that respect yes in order to get commercial entities to participate in a common task method do you think the sort of traditional evaluation and task definition would need to shift to include uh artifacts like scale speed adaptability to different domains that's always been true since the very beginning actually so uh uh the the very first tipster and track work funded by the defense department was about document retrieval at scale and i mean the scale in those days was different this was the 1980s late 1980s but by scale they meant let's let's try information retrieval not on 1500 documents but on a million documents and nobody had ever been in a position to do that before because nobody had a digital database of a million documents or at least hardly any people did and so they they put together such a collection uh similarly for speech recognition throughout the 1990s um in the competitions that the government ran there were always tracks that depended on speed there was a sort of you know real-time track and a fi less than five times real-time track and a less than and you know an unlimited time track i don't remember exactly the details but the idea of of uh seeing what you could do under constraints of under under resource constraints was uh was definitely part of the picture so what you say is absolutely true but it's also been true for a while i think yeah you should maybe share uh a couple more reasons or inclinations why you are optimistic that uh the issues that you are potentially alluding to like companies uh being more secretive about their methods and their algorithms and their data sets what are you seeing out there right now how do you encourage companies to participate in community settings in cattle competitions rather than doing secretive things and why you have this well so so why why am i optimistic about companies not becoming more secretive um and the answer is that i i didn't actually say that i was optimistic about companies not becoming more secretive what i said was that i'm optimistic about progress continuing one way or another and uh um i suspect it's the case that that many companies many companies already are quite secretive and i it wouldn't surprise me to find that more are more secretive because if you have a situation where a market is sort of you know as big as it's going to get or the basically saturated and the question is what percentage of it do you have compared to your competitors then the the the incentive for a company to try to uh improve the basic technology on a broad scale so as to make the market 10 times larger so that even if they only have the same frac you know companies got 40 of the market 40 of a 10 times larger market is very attractive even if there are lots of other people who are making money too so to the extent that there are areas where the technology is well not only where the technology is not yet mature but also where by making the technology better there's stuff you could do there's things you could sell there's you know technology there's products that could exist that can't exist now there i think that just the logic of the situation is that if there's a company that's well positioned to take advantage of that increase in the in the market it's actually to their advantage to get as many people as they can possibly recruit working together on the problem um so the fact that there are lots of problems like that makes me reasonably confident that there will continue to be some companies who respond to that logic in the logical way uh but i i would not be surprised to find that you know precisely because this is now an area where lots of people can now be making real money doing real things on a large scale that you would get more secretive stuff whether on the part of startups or on the part of established companies but i guess the other reason another reason that i'm cautiously optimistic is that an awful i mean we're still even more pr definitely more than 20 or 30 years ago we're in a world of tinkerers you know this is like the time when a couple of brothers in a bicycle shop in ohio could build an airplane you know you didn't have to have a billion dollars in a big plant and all kinds of fancy equipment to build an airplane because airplanes were really crude and there were lots of people who had enough machine tools and enough intelligence and enough initiative to do it and and now you know any kid with a laptop and access to the internet can explore all kinds of interesting problems and so you know it's it's a lot easier to do speech and nlp research now both from the point of view of how much it costs the startup costs for the resources in terms of the the uh data that's sort of freely available on the internet in terms of the software you can download and and so forth so i think that the the part of the community the part of the process that involves people who are not yet exactly commercially enmeshed is also kind of promising looking you