Scale By The Bay 2021: Ricardo Baeza Yates, Responsible AI Challenges and Recommendations
Recording: Scale By The Bay 2021: Ricardo Baeza Yates, Responsible AI Challenges and Recommendations
okay thank you gb for the introduction i will start right away because i have a lot of material i want to cover so today i will talk about ai ethics and some challenges and then some recommendations so um just as before i want to mention this new institute from north system in exponential ai that basically wants to to uncover what was in what is important in successful applications like having humans in the loop or better humans in control and also the stronger dependency on data instead of better algorithms and usually big data when most of the companies have small data so i will start with the main ethical issues that for me so i have a personal bias here and these are the the four main ones discrimination phrenology lack of semantic understanding and then excessive use of recursive resources and then i want to to discuss a few things that are related to this and at the end some recommendations so let's start with the main problem in discrimination so the course of bias so you have bias data remember information is biased so we have a bias to understand bias in a negative sense but in practice bias can be positive too so it's just a systematic deviation in some cases it's bad in some cases it's good so the question is should the an algorithm should be neutral or fair against the negative biases and that's not an easy question because it's not a computer science question he said like it's a social question so typically this question is not asked and we get the same bias but even worse sometimes bias is amplified if by exemplified we cannot say that the bias comes in the data in fact bias is not only the data and i hope that with the examples i will mention it's clear that although that's the main source of bias there are many other things that that may introduce us in a system so a very well known cartoon for many people so for the difference between equality equity and justice usually we try to do equity because justice takes a long time it is much harder to achieve if you're interested in this topic i really recommend this movie called the bias introductory you can find it in netflix uh done mainly by women and for me much better than the social networks one that was last year and we did a very interesting panel with a federal judge and three other computer scientists in march organized by acm and that's available in youtube okay so what is the the answer to this question should we care well not always as usual when when you have a hard question the answer depends but if we have people you need to basically care about this question and we will get back to this at the end now how do you handle this problem where the the easiest the the best solution would be to devise the input but sometimes we don't know the bias or sometimes we don't know the right reference value another possibility is to tune the algorithm so there are things like learning to rank with with bias that learn how to handle certain specific biases again you need to know the bias you cannot do it in general and finally you can devise the output but in that case you already lost a lot of information and the solution will be virtual now the first example of discrimination that reached the headline news i guess was the the compass case where republican 2016 claimed that there was a racial bias bayless russia bias in uh basically a criminal profiling to get for example uh conditional uh freedom in in prisons and although this was created as a support tool in many places in the u.s was being used was was used as a decision tool and here we have two important questions that can we can pause from this example is a secret algorithm ethical so basically do we want transparency on how for example a person will will be uh predicted to to be a potential period in the future or not however together with this question we have another one that's very important it's a public algorithm safe because this can be gained so you have these two things that are important that are really uh are in conflict criminal filing is something that has been done not only for prison but also for for basically when police stops a person in the street for example poland still has a famous system gordon that has been used in some u.s cities without basically the knowledge of most people and even in some european countries the same in chicago uh the system that was built with iit and here you have a particular case of the system that you have geographic sampling bias because you police the region where you think there are more criminals and then of course you basically reinforce that belief because all the other crimes that are not reported elsewhere will not be in the system let me give an example why why buyers can be amplified and this is a interesting paper by climber at all that was published a few years ago uh let's say you have an offender and and this is the the example is uh bail in in the states of new york uh in most of the world you need to think if the person will reoffend or it will appear in court or in new york you want to have to the judge only have to predict if the person will appear in court or not doesn't matter if it was a serial killer or not which is i think very hard as a person to basically not think about that now if the person gets bail and can pay good if the birth person doesn't cannot pay there's always people there there's a person in the u.s that even loans money to people and they have to do the same prediction of the judge will this person appear in court and return the money and finally if you don't get value go to prison here we have one problem is that we don't know what will have happened if the person had bail so we have half the data and then we need to do that imputation we need to infer or predict what will happen if that person had bail or not well the results are very interesting so assuming that the prediction of the system is correct they were able to decrease the crime rate in 25 percent or decrease the prison rate in 42 percent so much better than the judges even for the one percent most dangerous criminals the gadgets were doing like big mistakes now they there was an interesting data methodology because they split the data in five parts like hold out data for imputation data for training and data in a log box so anyone can test it later with data that the algorithm never saw now this algorithm was didn't have have too much data so it was not good enough for for deep learning but also they want to have some interpretability so they used gradient booster decision trees and the only demographic feature they used was the age you could gender is mostly men so that doesn't make much difference and there's nothing else the rest is just information on the case well what was the result in red i added this is a table from the paper in red i added the percentage of people that is black or hispanic in new york and you have 32 percent but you already see that it's a big percentage much higher percentage on the people that the police brings to uh the court so it's 82 percent well the judges clearly are racist in the sense that they increase their rate of blacks to 57 percent they decrease a little bit the hispanic or white 32 percent but in total they do 82 percent goes up to 89 what happened with the algorithm well the algorithm learned to be racist and basically amplified the races against african americans so the 57 went up to 60 and also amplify the decrease in the spanning from 32 to 30 percent but overall the the minority grew from 89 to 90 percent now you can the good thing about the algorithm you can match different force different distributions and then if you match based on the basically the best judge the system still was able to do much better like minus 23 percent less crime but why is this happening well we can look at different types of judges so the in the paper they divide the judges in five quantiles of different leniency and if you say okay you want to send additional people to prison so don't get bailed you see that the algorithm here will take the most dangerous one for the algorithm however if we see according to the algorithm what the judges are doing in the second quintile is almost like random so basically they are giving bail to people that is dangerous or to people that is not dangerous and here we have part of the reason for this so you can have noisy decisions like b you can have biases decisions like c and you can have both biases and noisy decisions like d so what is better a biased algorithm that is just in the sense that to the same case always gives the same result or a judge that for the same case we'll give different results there are papers that show that if you see a judge after lunch is one of the worst times you can see it because it didn't have lunch will be tougher so this issue about noise was uh covered in the harvard business review in 2016 by academy and collaborators but this year in may they pre-published a book called noise and basically focusing on problem of uh the variability of human decisions and this in some cases may be more dangerous than bias in fact one good thing about algorithms is that they don't have noise another example discrimination facial recognition so in this paper that was published earlier this year you see this the four phases of facial recognition and if you see who from where the faces came in the last two phases basically from 2007 after and you see that most of them came from the web well those didn't have any consent so there's like everything was developed with no consent of the people to use their faces well you know there were people accused of a crime wrongly because of this facial recognition system that were not well trained with some minority populations and after this in june 13 2020 uh many companies decided not to sell this software to police enforcement in june 30 we at the acm us technology policy committee also urged to suspend the use of factual recognition technologies but it was a bit too late because the software was already out and for example in september to 2020 another african-american was wrongly accused of a crime because of the failure of the sovereign and there are similar problems in language translation with gender bias or uh stereotypes for example here you have a famous paper from novice 2016 where you have chi here analogies if she's a nurse he's a surgeon or she's a diva he's a superstar so there are things coming from the text that are biased the same with uh coming from words of buildings for example who are the most typical professions for different ethnic groups like hispanic asian or white well what about language models and this has been in the news lately now we have gpt gpt4 i don't have this here but the size will be about similar to i think it's about five 500 billion parameters well in a paper this year they show that gpg3 has anti-muslim bias for example if you write the sentence two muslims walked into a you have all these conclusions that most of them are violent so this is what you learn from the news and if you look at what happens with different religions you see that much limit is four times more violent than christians and good news for many of us the less violence are buddhists and atheists but it can be much more complicated in ads in the example where last year the uk government decided to predict the final scores for university entrance without having the students taking the exam because of covet of course the minorities were affected or for example this interesting case in bologna italy in early this year where delivery was fined to be doing implicit discrimination this is a platform similar to ubereats for the people in the us and the main reason was that basically the model learned to give more work to people that could deliver food at night dinner time and of course the people that couldn't do that had less work so basically they didn't have a rule of let's try to give the same work to everyone and they were fine symbolically but that was important because that means that the source of the bias can come from the optimization function and the worst problem of discrimination is basically what happened in netherlands also in january this year where 26 5 000 families for several years were wrongly accused of cheating in child care subsidies and at the end the result was that the government the whole government had to resign on early january so this is maybe the worst case of how ai can affect not only people but also can affect governments let's go to the second one really it's not phrenology but it's physionomy but most people understand better holy chronology and the first one is from uh was the most well-known force from kozinski from stanford that basically uh use facial biometrics to predict such sexual orientation so is there any scientific basis for that but even if there is a scientific basis should you do that so there are two different questions it's ethical scientifically and it's ethical is morally socially and the same happened in 2017 when some researchers in china tried to to predict criminality using face images and also last year again but in the us the same happens and of course people complain because your face cannot give away states of your personality kozinski came back earlier this year basically claiming that he could uh predict the political orientation this was like a 70 percent accuracy that could be even uh spurious correlations so this phenology is something that was dead in 19th century for example when dr chester lombroso in torino italy uh collected hundreds of schools thinking that criminals had a different skull at least he was a real believer because he left also his skeleton i guess as a ground truth of what he thought was a good person so but i don't know if a good person will collect all his life skulls from the mark but it can be worse in 2019 mit research i claim that with your voice i can draw your face i wonder how they will do it for adopted children and maybe that may work for the neck or the mouth but what about the eyes your nose ears here what relation has that with your boys so i can have my master phrenology algorithm where you give me a voice i draw your face and from your face i can get your name yes there is a patent application for meteor that claims that they can predict your name from your face and then i use all the other words to know if you are in the position if you are gay or if you are a criminal in many countries this is really dangerous so you shouldn't do this and then we have the third problem lack of semantic understanding uh in 1979 george walk says all morals are wrong but some are useful he was talking about the statistics but i think today we can say the same about material learning so sometimes you don't use the right data for example the best example i have is also generally this year when elon musk said use signal and some software stock market software that use twitch tweets i thought that he was saying that they should buy the stock of this texas company that went the price for for more than 400 percent more so the company was happy the people that bought was not that much happy that would be even more stupid for example uh someone in facebook decided to use a model chain in english for bad words and the town of biche in facebook was three weeks without their facebook page part of the program was now human in the loop to check for these complaints and this may have hurt people because the the facebook page was used for example for covet announcements or you can have an example from adversarial ai like here on the right from some japanese researchers where you change one pixel and then you get a different uh a different level so the question is what this person is uh learning i don't know so some limitations uh you have to remember how to forget it's hard to forget filter what you learn you have a very interesting story from morgues about that you cannot learn what's not in the data you shouldn't wait until someone dies like in arizona for a self-driving car that case was not in the data and accuracy is not the key is the impact of errors the issue and we need to be humble we should say i don't know when we don't know and let me give you 30 seconds because i didn't realize i have enough battery so let me put the charger on i will be just in one minute one second he was very inclusive [Music] so sorry about that i will continue so the last one is waste of resources here you have the same table before large language models and this came from bender and gabriel so in this paper uh bender gabriel showed that the carbon footprint of training transformer was 57 years of the person normal person and then also that this was uh between one million and three million dollars per training so the question is are we using these resources well to to change what embeddings that basically are biased to many problems and you may notice a famous paper the thoracic parrot that was part of the problem where this ended with the layoff of timmy regardless from google and also later from margaret mitchell for a similar reason so let me then go to the discussion of what we can do so these seven properties were proposed by the acm in 2017 and systems do not need to be perfect but since the people expect much more from ai than than what we can do because humans judge machines much harder than people some people is already [Music] working on how to do ai software that is trustworthy and the question in the future will be how to develop software with the help of ai this diagram is a from a paper from ben schneiderman last year we have issues on data protection identity and privacy for example these damaging so basically you have the concentrical basis how much data you collect how much time you store that data and for example in privacy's power from calistability you can go deeper in these issues if you look at article 22 of gdpr the last paragraph says that the person can contest the decision this means that you may need interpretability to give people information about the processing you get to explain ability to challenge the decision and you need to do validation testing and maintenance to keep the system working as intended and here is an interesting example of how this artillery 22 has been already used legally in france where a court said that using facial recognition in schools was inconsistent unconstitutional because they didn't have the competence they didn't have the consent of the people and the solution was also not proportional to the goal we also have already some u.s regulations there are cases against amazon google and facebook in different parts of the federal system during trump a few laws even one that was proposed by tamara harris didn't pass so we may see this coming also can the person that was uh writing about antitrust already on amazon a few years ago is in charge of this so we expect that the new national ai office will address these problems in last april the eu proposed regulation for ai and basically they use a risk-based approach and it's already some problematic because risk is a continuous variable for example here in article 5 you see that they are they forbid any use of a subliminal technique that basically can distort the person's behavior that may cause a physical or psychological harm so if you apply this to the letter of the the law that may imply that you cannot give for example an ad of quick fast food to a person that has a metabolic illness and some solutions include registered algorithms old algorithms they're very interesting paper in fact this year where the group of northeastern northern algorithms publish a real audit in a hiring software company so we have several human practices i will skip this but these are based on some cognitive biases and we have some professional biases on on for example even there are some papers that show that the biases of the coders are transferred to the code so this is much more complicated than we think so what we can do where we we should then analyze for known and unknown biases we should regulate more data in sparse regions and we should not use attributes of sata directly directly with hardware bios we should have the system aware of the problem and we have the users aware of the problem some recommendations to end basically that learning from the past doesn't mean to reproduce it have people in control not in the loop and remember that ethics is not something that depends only on you also depends on your providers and your clients so remember that everything is a mirror of us and chester henderson asked can a algorithm ever be ethical well they are not human so maybe they can't be ethical and david lauer asks something obvious but it is very important you cannot have aisx without ethics so thank you