Devreal

SF Scala: Omar Alonso Interview

SF Scala: Omar Alonso Interview

Recording: SF Scala: Omar Alonso Interview

[Music] hello everybody I'm Alexey Krylov the father an organizer of by Eric this is our first meetup of 2019 we're here on location on Domino data labs which we've seen grow from a few people at galvanize the big company now they use Cullen the backhand and the serve data scientists in the cloud and so we're happy to be here on location our first speaker is a moral Lanza whose principal software engineer at Microsoft and he is a veteran by the bands he talked about data data by the bay before and so we're very happy to have it here thank you very much romantic saying so can you tell us a little bit if you know basically what your focuses and what have you been doing for last couple years sure so my interest is mostly on the intersection of information retrieval knowledge graphs and human computation and we were working obviously with a lot of social data in the last two years we've been doing a lot of knowledge graph extraction from different sources not only Twitter what also other Microsoft data sources the talk today is one of those projects but it's basically mainly about unsupervised generation of knowledge graphs mmm so are you I think you can't given talks about label quality at our previous event and I think that was one of the really foundational topics so I'm curious how did you come to the area you know like what what is it about data quality we should talk to you one of the reasons is if you want to play with new data sets like information retrieval and you want to work on search at some point you have to do evaluation and if you don't have ground truth it's very difficult to build those data sets and when I was at a nine I you know learn Mechanical Turk and I got hooked into the idea of doing labeling in general that's how I got to do labeling that's how by the way I gave the talk and thanks to you but I'm trying to finish a book thank selected for an invitation because that kind of we're trying to grow these slides together in a short monograph which is mostly about the practice of labeling mm-hmm there's a lot of publications on how to do labeling but the bulk is how to do this at scale and involves learning how to design your hits your human intelligent tasks and then based on the domain applying different techniques that's basically the the practice so this is I think you know when I think when you did this talk the whole like AI wave was not as fine as it is to me and now I think what should we see better relevant examples right because the trust in machine learning now I think more people understand and I think there were like mainstream media articles about bias right so if you know i strangle certain kind of people it will not understand all the kind of people their voices right like the characteristics and so forth so so I'm wondering kind of do you find that people understand that the talk comes down to the data originally like did you see more people asking you about this like : to amp the data quality oh boy yes so the answer is yes people are getting more interesting and how to how do I build training sets how do I assess the quality of my machine learning models ai models at the same time biases is and you think that is coming up a lot and last but not least the notion of annotations on labeling datasets because sometimes you don't know how you gather this data sets right so I don't even I was the sampling methodology which areas we touch etcetera etc is usually not even thought and it's it's been it's becoming an crucial because you want labels you want clean the reset and then you not only you want to know how buys the data studies but also you want to know all the annotations so if you want to reproduce another experiment so you know to refresh a data set you have to have all the different knobs that you used to build this data set which in the past was like God gives Alex a data set it was right that's good now it's like you have to derive these data sets over and over again for million different properties and domains so you know what comes to mind is rather let's say what's this called fake news phenomena no right obviously news people have different opinions so if you ask if you want to ask a bunch of people to label political tweets Russians will label them very different than Americans in Russian Americans will label them different from either of these groups yes right and so so I wonder do you find like they do try to compensate for this they want to understand what biases like humans have biases and like do you want to go back to them and understand what biases they could have had like how do we have to understand biases in humans like do we have to go back all the way and try to model I think what you need to you need to do if you're going to build a set of sets is to you get an idea if the task has a subjective answer an objective answer or somewhere in the middle so yes if the task is about building objective answers then you can get agreements and compute in traded coefficients etcetera there are certain things like election as Democrats versus Republicans you know fake news versus no fake news those tend to get much harder so it's difficult to get full agreement but what you want to know is if you sample different times the proportions are stable which means like well if 205 thing it's fake always maybe there's something there buddy one sample nobody thinks is fixed on the next one everyone thinks it's fake then you know there's something fishy there right so it's more about trying to find those tasks were things you're not going to get agreement on it and then do something with it versus blindly you know either either blaming the data set or the workers or the interactions which is usually what is done in practice understanding the nature of the problem fake news detection is hard mm-hmm because like I said if you think in terms of US vs. Russian politics is one thing but if you say faking news about a brexit maybe it's a whole different thing right maybe people make get more agreement on fake news in Rex's versus politics in the u.s. right so understanding that I think is kind of the precondition for doing the rest right and fake maybe I mean an objective through like something happened and news is wrong then you can say it's fake but if something is political opinion right it's conventionally a question because somebody speaking is can be blown into some other people I think it's super hard crack and also fake for you may have a different definition than for me maybe you know you find the news that is so fake that is hilarious right versus a lot of persons say in the midwives or in these girls will get offensive that's right so understanding the how humans will answer to the question is this fake or not it's also very important so you'd never have to take it for granted and say oh just go and grab some fake news data set because there's always gonna be very difficult yeah so you know it's so like that's the ground true but now you you are into knowledge broths yeah so how do you move from labels to large knowledge drops it's an internal project so I've been working on the reason why I got into labels is to leverage human sensing and with social data say Twitter or Facebook there's a lot of human computation at scale so why not just deriving graph from that sensing mm-hmm so if you've already stood in today instead of like counting hashtags and through the topic can you derive a graph if you have a lot of been sharing with veena social network can you derive something else so basically aggregations that you can do that are just beyond counting terms or topics but if you can derive these connections has a lot of labeling yes Raven is also hard because now you have to label the quality of the graphs but it's kind of like the next step on on labeling from my perspective in my own products I just say so knowledge graphs is an established area right and so on and if in my mind it was like the what comes to mind the meters are DF and Semantic Web yes and it's kind of it was cool for some time yes but not recently well I think it was there was a like it's a movie before I die before I kill mother I'm kind of fashionable I write there was a notion that like at some point it looked like the bill Semantic Web everybody's gonna very like painstakingly underneath their website protocol and we'll be semantic triples everywhere and didn't materialize like it's really I think the feeling like there are some like hardcore veterans who still do that right and I feel like I've seen a lot of work from German universities and I think I can see my catalog and everything but but where is this field I'd like and I met people from Google I think Eugene Gabriel of which was knowledge graphs right so but I think they didn't like it internally to do some modeling so but I don't see you know like mainstream companies talking a lot about it like what is it is this still a very like but I know it's used internally I just don't know how much with this proprietary how much of it is you know open research like what's the state of knowledge graphs great question so our approach to knowledge graph is not the traditional AI what you just mentioned it's basically more like domain-specific graphs which have some knowledge on it you mentioned Eugene who is at Google Google has has this Google knowledge graph Microsoft has the Satori knowledge graph another companies like Amazon they're willing in all knowledge graph for products Pinterest is also building a knowledge graph for red across the street so you can think of like little little graphs that are specific for the domains entertainment business etc so the goal here is not to capture the world view so that we cannot the psyche and all those projects but it's more like for a specific domain say music you know who are the black what the bands were the players were the genre etc and have all these connections if possible automatically derive mmm so it's not it's more about a data-driven approach to detect entities entity resolution and linkage etc then you can expose us a graph with annotation so you can color knowledge graph that's how we call it it's not like a super you know the world view of although or the planet but it's a good representation of a specific domain interesting so this comes from original questionnaires or is it derived automatically like what do we need from humans the levels of human rights in the case of the teskigi graph that will we'll talk today is about detecting links people topics detect that and make a few connections you can build some pre interesting applications you don't have to have you don't have to go full fledge into super crazy techniques which is understanding people places organizations links topics and then make the connections the occurrences or anything else and you can have a pre decent graph I say graph we could call it knowledge graph because have some knowledge in it but it's not RDF and we don't know we already care on the representation of it it's mostly like which are the different entry points you can choreograph and get data around it that before was difficult in the example of Twitter you can see what people say so for the u.s. elections two years ago before the u.s. elections which is something is difficult to see today in to her this is just my brother these little little artifacts you see interesting so it's kind of domain-specific knowledge drops domain-specific knowledge wrap perhaps is a better it's a better terminology is lightweight that let's put it this way lightweight cavies which are not super comprehensive I mean for the for at an encyclopedia like Wikipedia but it's very good at the domain unlike hierarchical or other flat people on topic so this is one level hierarchy of topics well which we feel so for the one that 1% we don't but if you if you have the baseline well done then aggregations are somewhat easy to develop and we'll show also a couple of examples on those aggregations which is that the the point of villanies domain-specific knowledge graph it's like the aggregation that you can build so it's not about the data information knowledge so that pyramid is mostly like if you have these little connections what can you do versus I don't have the connections therefore it's difficult to do this mmm-hmm interesting so I wasn't gonna shoot a little bit in sense the maybe connection maybe not but let's see so so I took to for a socialite who's the Kira's creator he's at Google and he's a Google brain right and I talked to him about deep learning beginning of this year last year which was like the peak of kind of deep yearning is gonna solve everything right so basically and he was very interested in the key he kind of was saying that deploring will never be general artificial intelligence in Mexico saying it will it's not capable of generalization because it's very good with pattern detection right but it will not generalize like if it give you different kind of some different kind of stripes if you give it a completely new kind of stripes I shall not you know the kind of so it will not understand it and so he was actually talking about that we need a fusion of statistical AI and logically I and I think through Russel Berger talked about that as well so basically that like the hinting at this like the limit of statistical AI which we are doing right now and they needs to be some combination with what you know cool used to be cool traditionally I right so we've got this knowledge basis or something else but basically like I was surprised to find you know among the kind of cutting-edge people and deploring the feeling that you know you need something else to go to the next level and and so I wonder like the work you do kind of is it kind of the kind of related can it be an automatic like building knowledge bases like basically we need world knowledge in some way we need something else and besides deep learning is this possible complement to deep learning or it's it's its own kind of specialized area do you see any interaction so for the record I don't know anything about deep learning so can't comment on that front there one of the reasons why we've done this bottom-up and supervised is going back to one of the other points you were mentioning is about explained ability so if you want to the tag that Alexia it is you know related to say dominoes so something how can you explain those relationships and sometimes a bottom-up approach and I supervise you can explain some of these things mm-hmm that's one of the reasons why the project that I'll show today or on Twitter is because we want to be able to explain why this thing is related to that and the other thing is provenance we want to show evidence that these two things are related and if you can sprinkle that they are set with provenance which is a well understood topic and databases then you can make sense of this pursuit of you know this thing is related because it's magic will be its ability because here's what we believe it's a relationship slightly maybe there's a connection I I don't know this one is bottom out that's what I'm trying to explain I supervise I'm from the ground up grassroots if you want to call it mhm kind of a lightweight KB's interesting yeah I'll let our viewers further comment when they see it cool so so well definite looking forward to talk about like overall where is your research going and like what are your plans for this year so for this year that was gonna be here in San Francisco mm-hmm in May I'm running with co-chairing we love the human computation track which we have some superb papers the conference is going to be massive so if you can attend please attend in October helping out with H comp the human computation covers near the border watching Tuesday the rest we have a few things or nons graphs that we would like to if we get a forward for Microsoft to publish that's it we mean like under under the covers in the trenches right to get you some products that's not a lot of not our research everything that I've mentioned today has been already published and the rest is well you know working hard to put it externally at the mall and it's great to hear because you know we in the community have for instance figure-eight for macro flour which do a lot of human yes can be dangerous I think this topic is actually kind of very interested into the motive of people in this community so this is great to have you here looking for their dog thank you for invitations [Music]