Devreal

LLM Avalanche: Panel: Trust

Event: LLM Avalanche SF 2023

LLM Avalanche: Panel: Trust

Recording: LLM Avalanche: Panel: Trust

foreign really a conversation because we are in the open field right when in The Big C we want to find our course we will do together so without further Ado some Galaxy crowber from one of the organizers of this I I run by area meet up and scale by the bay conference for 10 years in in the Bay Area and the tenth one is in November it will be like this but 10 times bigger so please join us in November the cfp is still open a few days more until end of June uh and I also run open source science program at IBM research and meet my colleagues at the open source science sign that was Paige talking about join us it's a volunteer effort to auxiliary science with open source so we welcome everyone in our letting everybody introduce themselves left or right start enough offer hi everyone I'm offer mendeleevich I run developer relations for Victora and kind of a way of introduction I had the opportunity to work on uh llm since it was gpt2 about four years ago so it's really interesting to see how it progressed over the last four years and kind of was very impressed with its capabilities and seen how but also some of his challenges uh and so I've joined vectara we do a platform to do to develop generative applications kind of seeing the potential there was a no-brainer for me to join a company and happy to have a lively conversation with everybody here thanks hi I'm Christopher Guin I'm the founder and CEO of company called itomatic and we work on Industrial Ai and and llms are at the heart of that and the reason that we discovered or the need for that was my previous company was acquired by Panasonic which as many of you may know is now a global Industrial company and I I you know there was a startup that I started a number of years ago many years ago in fact I was one of the earliest committers to to spark I was a machine Learning Company and my job at Panasonic I I ran Global industrial Ai and and we failed right very quickly and I I realized that what we learned what we knew of as digital AI or digital machine learning that we do in Silicon Valley actually does it really work for the physical world I'll be talking about that at a later talk but obviously the work that I do at Panasonic there's a lot of uh dependency on safety reliability and so laundry avionics automotive and and and so on so I've had a long career here but uh you know this is uh this is a really exciting times I like to tell my team that this is not even a one generation kind of Technology disruption probably once in a millennia so looking forward to talking about that hi my name is Sandeep gopesetti I work for IBM research my career started with handwriting recognition rights and I used to work very closely with the media you know with images with speech and since then you know the world's changed quite a bit right in terms of the use and the adoption of AI with the introduction of chat GPT and the GPT models in particular you know it's been everything for everyone all at once so uh the key challenge now is that how do we trust it right and hopefully we'll get to delve a little bit more into it today and my new role is really about data and data management and the use of data and for Enterprises so my focus is purely in providing better AI solutions for Enterprises hey everyone my name is Josh Tobin I'm the co-founder and CEO of Gantry we provide infrastructure and sort of workflow oriented tools to help companies with all of the stuff that happens after you deploy a model so really focus on maintaining and continually improving models in production and one of the key questions that our customers always have there is hey how do we make sure that these models that we're developing or that we're trying to improve are actually trustworthy and they're actually safe so super excited about the topic today I also co-organize a course called Full stack deep learning where we were sort of one of the first classes to teach folks about all the stuff that you need to do to make models work as part of products other than training them which turns out is you know sadly for myself as a former machine learning researcher not the most important part and previously I was a research scientist at open AI hi everyone my name is changing I'm a senior data science consultant at databricks I help customers Implement scalable machine Learning Systems my journey at lrm started when I was in grad school when Sesame characters were all the rich you know the bird and another elbow models and I'm also one of a co-instructor for the edx course for large language models in production thank you everybody for your introduction uh just remember we can give a mic to uh changing for now right and then give it to the audience so everybody has a mic thank you all right so trust right this is a big topic I think this is the key and crucial topic and the whole uh area right now and uh a lot of you are actually uh experienced I think some of the most experienced people in the world right now because the whole thing started right a few months ago and so uh uh and obviously we can come uh uh at this from different angles I'll ask you to very quickly say what is your definition of trust everybody Define trust in like the shortest way possible and maybe link it to your to your topic to what you do yeah I'm happy to start uh so for me trust uh in the context of lrms is a couple of things one is I want to trust that the model is going to give me the right answer right so we we know about this for a long time there's this sorry not a long time six months but uh there's hallucinations uh that's a big problem right and other things like that so give me the right uh the right answer trust is also kind of another angle I think about is uh privacy of my data how can I trust that the data I use when asking questions or when using my own data with large language models is not leaking there was a big a big thing about that and so that's the other angle I'm thinking about hi this is so the way I think of it is actually I asked the question why are we talking about it all right um on the one hand if you think about the kind of work that companies like Panasonic does it's a over 100 years old uh it built for example it makes all the batteries for Tesla at least the ones that are outside of China it deals with avionics Automotive refrigeration systems so clearly trust reliability is really an inherent part of their product process in fact there's an organization large organization in a company like that in the case of Panasonic it's called Product security Center and it deals with all of these issues um so in that sense it's always been inherent but on the other hand to a lot of us in Silicon Valley we've been you know I I launched Gmail at Google and you know the word trust at least never explicitly came up as as a major concern so I think the reason you know trust you know comes to the fore is is both a technological phenomenon but also a social sociological phenomenon as far as silicon what I use the third Silicon Valley it's sort of a proxy for for all of us here right wherever wherever we are um it's it's a new phenomenon to us where for the first time we are building machines that are actually they seem to be in the position of telling us what to do as opposed to be the other way around and I think you know when you think about trust in terms of trusting AI there's a lot of concern about hallucination uh if you why do we care we care because we're saying hey can can we really listen to these machines right so I think in that sense it is indeed sort of a new thing but I think we can draw a lot of lessons from companies like Panasonic IBM apple and so on in terms you know when we think about AI Trust both Oprah and Chris talked about hallucinations I'll try and not touch that so when we're talking about AI systems you know can be trusted not to be abusive right can we trusted not to be profane can we trusted not to have hate in it right in this day and age when pretty much you know profanity hate and abuse has been normalized in the society how do we get machine Learning Systems not to be abusive right how do you eliminate any of these things how can we trust it not to um you know be profane and abusive to the users the other thing when we talk about trust is you know personal information you know especially if it can be identifiable to a particular person right going beyond the wikipedias of the world or the papers or the Publications that are typically there you know how do we trust AI assistance to not make up stuff right um and the significant part of that is being able to teach a large language model about personally identifiable information and then understand what is objectionable in that right and uh significant amount of the problems when we're talking about large language models comes from the fact that there is redundancy in the way people interpret data and how do we eliminate that in particular right and uh in the good old days you know I mean if there was ever a doubt about a particular pattern and the label associated with it you try to get more people to label it right in this day and age when we're dealing with petabytes of data how do we really get uh you know people to uh you know rely on that particular data in a more trustworthy form for me trust comes in different forms not just in the you know the in the output of the llm but also being able to address the output of the llm and all the way to the data I think um for those of us in Silicon Valley it's um we have a tendency to get over excited about new technologies um and llms are maybe an extreme example of that and the reason why people are so excited about them is because of just the obvious capability that this technology has but at the end of the day you know 50 years from now 100 years from now when people look back on this technology they're not going to judge it based on how capable it was they're going to judge it based on the impact that it ended up having on the world and so the the what's going to lead to this technology being successful or not is how um widely and how quickly it's adopted in actual applications that solve real problems for people in the world um and so trust is the ability for the people that are involved in building deploying and using these applications that are powered by language models it's it's um it's the ability for them to take something that is capable of solving their problem and actually let it solve their problem um and so the main thing that's I think really holding back llms um there would is consistent with what's been holding back machine learning as a field um for you know at least the past 10 years which is that this is a technology that um inherently always makes mistakes right like if when we build software there's bugs and software and we're kind of used to that but people see bugs as these sort of rare things these uh these sort of artifacts of our imperfection as humans in developing this application but something that we can strive to not have machine learning is inherently a probabilistic technology machine learning models will always make mistakes and so whether it's hallucinations that everyone's talking about today or um you know some future problem that people have Downstream of that there's always going to be failures of these types of models at least until you know until we get uh well past AGI um and so the the real I think key to addressing the trust Gap is in my mind there's sort of two pieces to it one is user experience which is kind of the user-facing side of the trust equation and then the other is measurement which is the sort of developer and deployer facing side of the trust equation so those I think are going to be the keys and as a heads-on implementation consultant the top question that I get asked often is how do I make sure that this model can be reproducible in a production system and I think that's something that's really surprising about our aims today even for Chad GPT I say if you were to pass in exactly the same sentiment and ask check your video classify you know if the sentiment you know is positive or negative or give it a rate a rating you know from 0 to 10 we might get a very different response like maybe for the same response or same prompt we get 7.5 versus 9 sometimes even a five so to me that's really surprising that we are putting trust in these are our systems uh in the sub lrms where we cannot even guarantee consistency in the output so even aside from hallucination I think there is also a big gap in how we can how can we trust these systems when we cannot even be sure that given the same input the output can be different thank you so this is I think a very comprehensive you know ideas of trust but I know that each of you is building something and addressing in reality so can can you tell in a few words what part of what you are building is really resonating with customers right what do people see from what you're telling them from what they see which they really like and feel like this is increasing Trust and the viewers everybody and then you wonder maybe from that section yeah let me go first and change the order yeah exactly just switch it up make it more random like a large language model yeah okay so um so what is the problems that we're really dealing with today uh has to do with the fact that you know there's just different types of labeling on the data itself right and uh we're also expecting the llms to actually um you know understand reasoning right and to try and Define reasoning into a large language model uh is a very difficult problem right I mean you it's reasoning is not something that you learn from massive amounts of text I mean that's the mathematical logic that goes behind it is very very different so when we're talking about consistency in results um you know most of the problem with the consistency of the results has to do with the fact that different people have different opinions about the same sentiment right and it may mean very different for the guys on the right versus the guys on the left right um in a very abstract way and when you have people who are very polarized you're bound to teach llns to be very you know polarized as well and that makes it very very difficult to actually get to one single answer at any given time right the um uh one of the challenging things about trust is it's a long tail um lagging indicator if you think about how to measure trust right so a lot of companies that we talk to are having this experience of hey um we got access to chat GPT or gpt4 we've built this product feature we shipped it took us like three or four weeks our users loved it you know um meaning we moved the metrics for us um and now they're trying to figure out like okay it we have this sort of one-time spike in usage but is this thing actually is this truly solving a problem for our users is this something that is that is going to cause us to retain more users um to you know uh uh to grow our user base over time and one of the challenges with that is that like a lot of people are using these tools now because they're shiny people are using like everyone wants to use the latest AI tools um but I think what's beneath the surface is that there's a sort of looming wave of churn for a lot of these tools because the um at the end of the day you might love your uh your question answering search tool not to pick on any of those companies because I think they're great but um if you if you use one of those things and you know eight times out of ten it gives you an answer that you can't rely on or you have to go check your work with Google it's not going to become ingrained in your behavior right so you'll use it a lot at first because it feels better but if you don't trust it you're not going to stick with the tool um and so I think the um the key to trust is you know again comes back to um can you measure is there a way for you to measure what is causing users to stick with this thing or not um so the thing that we're building that's really resonating with with companies that are building tools like this is building evaluations of models that are tied to things that happen in the end user experience so if if you see end your users submitting particular types of inputs and then churning or you know being less likely to come back then you want to make that part of your evaluation Suite so that you can measure progress against whether um you're you're improving your model's capabilities on those types of tasks lining behind it's really that we as a community have come to realize how limited the traditional evaluation metrics are you know we like to rely on Benchmark competitions we like to say oh this is the top model in this leaderboard but it will now we are quickly realizing because based on the really accessible output that hey this is actually really subjective how who are you to say that this is the top response versus you know the the Das response so I think we are starting to finally get the subjective nuances in what large language model can output you know even though language models have been around for a really long time so uh yeah so for us at Victoria one of the things that we tackle of course is hallucination uh that's why I speak about it but uh I think it's interesting what we do is we call it grounder generation which is a form of a retrieval augmented generation it's a particular type of application that's pretty Broad and how it works is you help the llm gain other facts using your own data for example uh using a very strong retrieval engine so think about it as if you give the LM facts that and you emphasize those facts and you say please respond based on these facts it gives us a higher chance of not hallucinating and actually knowing what the base effects on and of course you get a response and you get the citations you can check but it definitely reduces hallucination significantly that's why our customers like it and and feel confident it's a it's an improvement on kind of just using the LM purely I think the other thing that we find uh very interesting a lot of interest from our customers is the fact that we because we only do that we don't need to and therefore can promise them we're never going to kind of use their data for training we train our retrieval engine which is different we train it on only public information so that gives people a lot of confidence as well I'll uh I'll talk about Alexa's question was what are your customers excited about or interested about when it comes to you know what what you do and as it relates to trust I'll try to answer that question by asking the audience a couple of questions and then we'll walk through it um one of course I think you all know Stanford Hai um does this annual survey of uh you know the state of AI and this year they publish something if you follow me on Twitter you'll see the answer to this question um but uh um so so they did they they published a survey done by by someone else uh essentially the optimism around AI as technology across different countries so I was curious and I took that and I plotted against GDP per capita right and so I the correlation is extremely strong the r squared is 67 percent okay so so we clearly there's a correlation but I want to ask you you know to guess like you know there's two polarities right one is richer countries are more optimistic or poorer countries or you know A and B is poorer countries you know emerging economies and so on are more optimistic in their response to you know optimism how many people think it's a how many people think it's B yeah you're so smart it certainly surprised me because the same data when asked about technology optimism it's exactly the opposite so there's something there we don't know what the cause is um the second thing is I in my work I deal a lot with sort of the the cxo level and also at the at the governmental level and so on um um and because I do industrial AI it turns out a lot of my customers are not in the US we have some in the US but a lot in Asia and Europe and so on where they're still making things right uh well we're trying to bring that back you know with inflation reduction not going to chips act and so on um but uh let me ask you and this is kind of a loaded question and I'm sorry if you work for some of the companies that I mentioned but when it comes to llms what degree of trust do you think these companies and these countries have for the sort of basing the foundation models on open AI or on Microsoft do you think it's you know it's let's just say it's High versus very low how many people think it's high low yeah it is very low and and it's not it has become a geopolitical uh you know Japan has declared it to be a national security issue not to be dependent on someone else's Foundation models so um I I mentioned earlier there's an interesting moment here you know the pastel analysis is not just technological there's a bunch of other letters political economic social environmental and so on uh legal but it's a really interesting moment where countries and certain companies say we don't want you know we want to we want our own models right so what what what we do at automatic we actually do something called SSM right you can guess it's sort of the opposite of llm there's we call them small specialist models and these are domain expert duplicates right and they collaborate in teams to solve very specific problems and of course the front end is natural language but you don't need all of the a compressor expert there's no need to know who the president of the United States was in 1875 so so based on that we can have models that are a thousand times smaller and a thousand times faster but also there are fully owned by by the customers and so that's what they're excited about they're excited about independence thank you this is I think this is a very comprehensive view of different aspects of trust I think we have about five minutes for us and I'd like to leave 10 minutes to the audience so you guys can prepare your questions so in this five minutes I want to kind of invert the typical panel Dynamics where you talk about your stuff so you heard about other people's stuff right and we're in the open source world if this is a community Meetup so given what you heard from others right I think everybody represents a company which is in business and so there is this term which I learned when I worked with Bosch which is an industrial intelligence company is you know as Chris describes so they invented some competition right because it's a competition on cooperation because all these industrial devices need to be working with each other and obviously a lot of pieces of software we use comprise a stack and an open source we actually explicitly use it so given what you heard like given where you are in the stack what would you like others to do to work with you better oh like if you can think of something and can anyone I can talk about offer right away I'm very excited about what Victoria does right because as you can imagine what we do is one level above that we use we definitely use retrieval augmented um essentially inferences right and so uh it turns out again once you think about the Paradigm of SSM small specialist models versus llm you don't really care about a lot of the intelligence in the in the the language model itself in fact the language model is really to us just a communication layer so sometimes I say I'm excited about this revolution in terms of communication more so than in terms of intelligence the intelligence can be provided by what Victora does which is backed by you know maybe rdbms maybe and PDS files and so on a domain specific knowledge that can somehow be mixed together by a very thin layer of of sort of a natural language translator well I should then do the opposite for you because I'm actually no no but I mean people you mentioned that SSM versus LM so I when I say llm I just mean language models and it's it's a true point right like we don't need it to be large they became large of necessity of you know they need to be more accurate so as if you looked two years ago you'll see like there's all these graphs of like how many model parameters you have and uh you know it reminded me of the old search days when Google and Yahoo and Microsoft thought of how many how many pages they had uh in their Index right so uh I think we're finding out in the research now that uh you know maybe they'll be larger or maybe not there's other techniques that are being held so I think a language we're agnostic to which language model we use underneath and and ssms could be another option for that as well so I'm excited that you're creating these specialized ones with your customers and um and I think again they don't need to be large they need to be like you said accurate and um we will add the custom data to them in whatever form they come so so I work in IBM research so I'm going to give a reset spin to this one right uh in the good old days I mean we had specification languages and most of the money was in the programmers who understood specification languages and programmed the systems to perform in a certain way um you know to you know to basically optimize whatever tasks that they were doing so now with large language models it's kind of flipped so what we really need are verification languages right so we need to have verification models you know how do we trust something that's coming out of a large language model how do you verify it and how do you really uh you know propagate that doesn't matter uh you know which industry you're applying to whether the problem is large or small uh I see uh a big need of investment and the verification phase uh to piggyback on that I I think we need to think Beyond just a large language model we need to think about a system that has a large language model so today I just finished teaching day day one of uh the data is Summit course on large language models and a lot of questions were about how do I make sure there's no prom hacking how do I make sure that there is a proper rate limiting how do I make sure there's no user abuse and all those questions have to come around like the infrastructure that supports this uh the single llm and it's very possible that you need more than one model that could be an LM or not to make this element actually usable and trustworthy in a production system maybe you need to have some processing uh post-processing system afterwards to correct your response before it even gets shown to the user so I think we need to just like how in ml in a machine learning context you know we think about machine learning model but I think we have fast forwarded you know to today we are finally talking about ml Ops thinking about it as a system and I think we also need to do the same thing for llms as well yeah I'll piggyback on the comments on um uh evaluation models or verification models since that's something that we're working on as well um I think one of the the keys is going to be okay so why do we need these things um we need these things because uh telling whether an llm got the answer right is difficult like we don't have good quantitative ways to say this summarization of this document is better than that summarization or um you know uh this um this generated marketing email is better than this other generated marketing email so we need intelligent systems to help us verify whether the things that we're producing as a result of these models are correct or at least trending in the right direction and um so the challenge that that presents is if you have llms evaluating other llms you might be asking yourself well who's watching The Watchman right like how do we know that these evaluation models are actually doing the right thing or not um and the answer to that is that at the end of the day um there's there's a certain like property of this this uh this generated text that we really care about maybe it's how much do humans like it or um how much do humans click on it or how much does it solve their problem and so I think the way that we all can contribute to moving this process forward is to um is to think of our part of our role as helping build and verify the the validation and evaluation models right so um the role that humans play in this is we are helping the developers of these systems gain confidence that their automated evaluation corresponds with the thing that we as humans really care about so you know give your uh give your friendly chatia a thumbs up or a thumbs down today and uh you know help progress of AI moving forward I think this was a great program so for now I want one of the guys to give up the mic to the audience if one of uh the guys can volunteer you know I think changing can keep yours I mean yeah you keep yours we have four guys one of them can give a mic I think all right thank you and um now please ask a question Jasmine will bring the mic to you the question is the front row and we have about five minutes so please be quick uh hi so when you discussed about the trust my name is lalit Pharma by the way so when you discussed about the trust uh I was hoping to hear something more from the so there are two aspects one is a training site and one is basically the verification side which you talked about I believe human in the loop is still the answer but in training side because some of the models are you know open source some of the models are proprietary how do you and something is evaluated by an l u not NLP alone right so it's natural understanding um what's your take on basically the fact that you know the models are so Arcane so uh proprietary kind of thing that kind of does not Inspire the trust yeah um there are two parts to it right so um one is that you know since most of these large language models are built on top of open you know openly available data and continuing to learn from a lot of these data as you generate go from one trillion tokens to two trillion to 5 trillion tokens um there is no way of actually preventing poisoning of language models right so that's in the training phase the second part of it is that you know when if you're fine-tuning any of the large language models um that's another place where potentially people could uh you know poison the language models themselves um when we're dealing with open source models right and uh you know from an Enterprise point of view or from an end user point of view you want to start building in some of these capabilities post um the inference right once it's in for instance you know you have some of the answers you want to run through you know potentially uh analytics that will help reduce you know eliminate bias eliminate uh you know some of the standard problems that you see with the hallucinations problems with hate abuse of profanity or pii right any of this stuff so you can start adding those modules rather than just accepting a large language model as is as your final solution right so it's going to be challenging in terms of how Enterprise systems are going to be built and how they're going to be used in the marketplace but that's where I see as the next big thing let's have another question for you so that another panelist can answer please so since uh for for since the data privacy is important and the data how you collect the trainer data to create a model is important so have you considered I mean what do you think about the federal learning to train the model without moving the data keep data private um I'm going to quickly answer that one and I'll pass it on this one of these people right so when we're talking about Federated learning um you know again it depends upon if you're trying to create an SSM or if you're trying to do take an llm and then customize it for the end user right I mean you know um typically in a Federated learning model you know you're limited to the amount of data that you can actually get from that one so you could potentially use uh you know Federated learning for um uh you know fine-tuning an application right not necessarily in creating a large language model itself unless you know your bandwidth and your compute and the amount of data that you have uh really exceeds uh the current limitations that we see there's a company called uh Laura cyber that that does um secure machine learning and they use homomorphic encryption you know and if of course they've worked a lot of the efficiency of it so that it's not like a thousand times slower um so so you can think look at Technologies like that on the other hand there's also again I talk about sort of this legislation to Europe that says for you to put a model into production you must publish the data that went into you know into the creation of it so uh of course that's for public consumption and so on so I think there's a tension between what technologies you want to use you know to say solve these privacy or security problems versus what the what what the concerns of societies are in terms of where these these models go I think the answer is going to be a collection of these things not any single one I think you earlier you mentioned that we should be thinking about this as a systems problem rather than just a model problem right the stuff that goes into it how it makes these decisions and the model interacting with other models you know in in a problem-solving Loop and the output of that uh you know I come back to my comment earlier there there's an entire industry that has been dealing with reliability trust you know safety issues for a very long time I think we can draw a lot of lessons it's just that our our internet if you will companies have not been organized that way and now we're talking about this and say okay we need to put that organization in place and you know you look at at the 2012 election uh Facebook Google and so on I've have to put some of these organizations in place to sort of uh you know try to combat essentially distrust so I think for companies to think about these issues at this time I think I think that's a good thing and it's going to be a collection of both technological as well as various other factors that that need to be brought to bear I think I think we really have only you know two minutes left so if anybody can summarize what we should do next in literally 30 seconds uh that would be good closer to us so your directive for all of us in 30 seconds what we should do next well uh for me I think uh you know I talk a lot about the hallucination I actually believe that this will get solved in a couple of different ways one of them will be retrieval augmented generation but I think if we think about gpt5 or the equivalent in any other company I think you'll get better there's a lot of research that happens in that space um I think uh I'm also interested to see how the regulation around the world will happen and what will it demand in terms of of that and especially I always mentioned open source models earlier I'm really excited about open source model I think that's I used to work at hortonworks which is an open source software company and now um seeing that happen in a in a space of larger it's very exciting I think you'll have a lot of benefit for everyone thank you for Chris and I'm in 30 seconds we're open sourcing SSM if you come to my talk I'll later I'll announce that there thank you um so from my perspective at least how long is it going to be before we start using generated data to start training large language models this is potentially going to be a big problem so identification of what is generated content would significantly help in advancing the state of large language models if you're building a product that's powered by large language models then think really carefully about the user experience that you can create that helps your users deal with the fact that your model is going to get things wrong um and there's not a technology fixed that's going to save us from that it's going to improve things but it's not going to eliminate them and then second if you're building applications with these things then think really carefully about measurements um you know just be like even if someone gives you access to all the data that gpt4 was trained on that is not going to help you trust gpt4 anymore because that is way more data than you're going to be able to do anything interesting with like that's think think about like trying to understand all the data on the internet um the key for these things is going to be um developing benchmarks that um intelligently surface like aggregate level insights about the model's performance on this massive amounts of data and that's something that closed Source companies are going to be able to do just as well as open source ones I would say there are two things first don't underestimate the power of your data so make sure you actually address that early on because we have seen the data problem trickling through the basically the age of machine learning to now and today we're still talking about quality of data and the second thing is tell discuss with your friends and even non-machine learning friends too you know about Chaturbate and the consequences I think the more people know about consequences of llms and what risks olympiations they have then the more power we have as Grassroots community in order to demand actually audits of models before they actually get released the public for use thank you everybody that was a great panel let's give a lot of those infamous [Applause]