LLM Avalanche: Panel: Build & Risks
Recording: LLM Avalanche: Panel: Build & Risks
foreign as I did it uh in the last panel if you all saw it I do not have the um the ability to give all the accolades that all of these incredibly hard hitters have on stage so I'm gonna let them go ahead and introduce themselves today as you can see we're going to be talking about for this last panel that's going to bring us on home and then take us home after this we're going to be talking about building risks and I would like ideally for you all to come away from this at the end with an idea of what you should be thinking about when it comes to risks there's many different ways that we can look at risks it could be engineering risk it could be looking at different data risks and then when you are building trying to keep those risks at the Forefront of your mind and recognizing that it's not after the fact that you want to be thinking about these different risks that we're talking about so we've got an incredible group of panelists I am going to start over here at Harrison and then just bring it on over go ahead and tell us who you are Mr Chase hello I'm Harrison uh CEO and co-founder of link chain which is a developer framework for building olm applications I'm your own I'm a CEO of robust intelligence a company that secures AI Bo OT and I work on scale and efficiency for language and vision oh and I I was leading Palm my name is Benjamin Harvey I'm the founder and CEO of AI squared I'm also the director of Designing trustworthy AI systems program at George Washington University wow all right and I'm Demetrius so this is going to be an awesome panel I can tell already when it comes to risks as I mentioned before there are a whole lot of them that we can look at so akancha maybe you can just give us a bit of an idea of when you think about risk what do you think of so when we think about risks we it's a broad definition and it falls in line with some of the AI principles that we have but basically starting with fairness and uh representation in the data that is used to train the models and then going from there to like training models and and testing the outputs for all kind of benchmarks and safety so there's a lot of testing that goes into the model outputs even before rather rlhf to make sure that these models are uh well uh are constructed properly and then there is uh of course application specific testing that has to happen and every application and downstream application might have very different requirements so that testing is very separate so that's where I would start the conversation also this is a ongoing field so as we will discuss further when these models are part of systems where you are chaining them perhaps with memory and other applications then there are further risks that you have to take into account including like are you running insecure code or are you doing any data leakage so all of those questions uh become even more pertinent based on which systems are interacting with the model sir Harrison and yaron I see you shaking your heads or not in your heads yes anything else you want to add on to this yeah I I think um I I I I wholeheartedly agree um I think that the way that we think about risk is like we think that there are three categories of risk um there is operational risk there's ethical risk there's security and you know privacy risks so we think about those in three different categories and the reason you should think about these or it's convenient to think about is in three different categories because they also sort of tell you who in the organization is responsible for them so you know if we were meeting here like two years ago we would mainly talk about operational risk that's the the one I think that like people in ml um you know in engineering would typically think about you know so these are things where you know traditionally you would be looking at you know where where could my model be failing due to kind of like operational reasons where it's like unseen categorical missing feature these types of things uh the ethical risk is like things that we started becoming I think as a community more aware of maybe I want to say like a year a year and a half ago when we're starting to think about bias right in in models and whatnot right and then that goes into like different parts of the organization typically Parts in the organization that would care about compliance that would hear about legal right um those parts and then last uh I think not least is security right um and that is you know for I think uh you know something that is somewhere in between ML and security uh that's nearer to the security Community but now the security community is actively thinking about you know about AI risk and that's again a different part of the different part of the organization and that's like you know those are things all the way from uh pickle file security um you know software supply chain aspects of it all the way to you know prompt injections adversarial attacks these kinds of things um so so those are kind of like I think the main three categories that one can think of yeah I think the the types of risks that I think are newer as was alluded to earlier when you start chaining a lot of these models up with real world systems and you have them not only outputting text but you're then using that text in a very specific way to do things with tools whether that's run code execute SQL or stuff like that and I think uh yeah to me that's uh very new and but that's probably I don't know to me that's the type of risk that I'm most interested in and also think can have a really big bad impact when you talk about in like if you think about AGI going wrong right a lot of it's it's doing stuff and a lot of that is like applying the outputs of whatever its model is to like a real world action and I think that's where a lot of the risk is introduced and we're obviously very early on in that Journey still but I think that's kind of like to me a lot of the new exciting and very important type of risk to talk about awesome yeah yeah go for it uh just to add to that the in the action space when we are using the models to take actions in the real world a lot of the real world deployments would rather go for scoring outputs which constrains the set of actions that the AI can take as opposed to open form generation and that's one risk mitigation that people look at all right cool so Benjamin when it comes to foundational models in contrast to other AI systems how do you think about the risks and um they're specifically you know these these Foundation models which are specifically trained and then used for a particular purpose what kind of risks are you looking at there great question so um one of the things that you'll see and I think the panelists up here have mentioned is that um at different stages of the machine learning life cycle there are different risks that could be introduced and one of the things that's important that that really separates kind of the traditional machine learning life cycle where you're a data scientist building a model that might be used and delivered to one or three different analysts inside of an organization versus a very general large language model is that as we are training these different types of models inside of these organizations these models are being delivered in a way where the analysts are interfacing with them from a human machine teaming perspective and because of that it allows us to have the ability to start identifying certain mechanisms of feedback so that we could start to you know understand you know the actionability of the model you know the relevancy of the model things like you know how timely the information is actually arriving and overall in an organization it's really about trust that leads to adoption right right so from an end user perspective you're really looking at how can I increase the amount of trust to overall for overall increase in adoption and that is you know really a component of you know all those different areas of you know can I make the results more actionable can I make them more relevant can I make it more timely can I add additional contextualization that can increase the trust and increase that overall adoption hmm I love that all right so it is uh for me very clear that when people think about different use cases with large language models right now there are two gigantic questions that come to mind and one we addressed in the last panel that I did and it was all about Roi and in this panel it's very much about the risks and the trust and when things go wrong how they can go wrong and uh I would be remiss to not mention my favorite part of large language models which are hallucinations and whoever was the researcher that came up with that name props to you because I can make jokes all day about that yeah like for example who hallucinates more Chachi PT or your average college student what I don't know yet so when it comes to hallucinations how do you think about mitigating those risks because it could seem really real and then you could put it into a commercial and put it out for the world to see that like this actually is happening or you know like oh you can you can damage your brand quite heavily if you are not aware of what is going on with these hallucinations so do any of you have any ideas on how to mitigate that and what you look at so first of all it's a it's a great point to keep in mind when you're building applications that these models hallucinate and they're actually overconfident in their answers so they will hallucinate with so much confidence that even if you're not an expert in an area you might read that text and you would be like oh that seems so believable I totally buy that so I think for information seeking they are not quite there yet it's an active area of research where some of the techniques include calibrating the models so that they can answer how they're how confident they are in addition to just spitting out text um there are other techniques which involve retrieval augmentation where we put the right set of text in the context window so those are some of the techniques for mitigation but in terms of deployment uh what's most useful is to have human in the loop who is an expert and these models mostly acting as an assistance for generating text or or whatever it is like whether it's code but there should be a human in the loop at this point in time to review what's being output and they better be an expert and someone who can critique the moral output is how I look at it awesome yarn I see you I see you wanting to go for it no no what do you got what do you got for us no I think um you know one of the um one of the things that you you know you learn uh when you when you hang a little bit with the security Community is that you learn that they have um you know one of the set of the controls that they have is just don't use the product right which is you know that's amazing right that's a very powerful control right um and you know that's not a bad idea right like for something so I think like um you know what um on one extreme you know I think that there is something about recognizing the fact that you know we're using models these models can you know hallucinate by Design almost so um we need to be thoughtful about the you know the applications that we're using them for and then you know sort of well if my application cannot tolerate um you know hallucinations maybe I shouldn't use it for that application I think that's that's the one to stream it but I think it's also you know like a very acceptable solution uh you know in many cases I think you know then there's a spectrum right right um all between you know somewhere between kind of just accepting the fact that their hallucinations were not you know we're just not accepting them at all and not using the model but then you know there's a spectrum in between and and I think that really um one needs to kind of the best mechanisms to try to mitigate those is to try to be really focused and not ask like you know how do I make sure that my model doesn't hallucinate but more how do I make sure that um for the specific tasks that I care about the model getting right um how do we you know try to design like kind of very well articulated validation you know rules or you know firewalls or whatever that you know that really nail down the you know the validations of the things that we care about so either don't use it you know at all uh for the if you have sensitive applications that you're worried about hallucinations or formulate um the problem a little bit better be more precise about the things that you really really care about and then build the right mechanisms and controls for those very specific things and what do those mechanisms look like what are you seeing as far as like the controls um again I think of it like you know you can um you can imagine kind of like an application for like a you know maybe you have a commercial commercial application and you really care about making sure that the you know now the model doesn't hallucinate about price right or it doesn't hallucinate about um you know maybe about URLs right and then it becomes you know from like this very big you know broad problem of like hallucinations to solving a very specific problem and that very specific problem that you're then trying to solve as you're trying to do like a lookup right almost to sort of see whether that URL actually exists or not right and that's a much easier problem to solve than you know to sort of you know go on you know this um uh heroic you know kind of uh fight against uh you know AI hallucinations I like that so uh Harrison I wanted to ask you I feel like there is a generative AI apps stack that's forming and I want to ask you because it feels like language chain is smack in the middle of it and when you think about okay if there is this stack that's forming what in your eyes are the building blocks as you're trying to build your generative AI use case yeah I mean I think it depends a little bit on the application obviously but I think at the core of it you've got the language model um and so I think there's a lot of uh choices there um you know open AI is probably the one that people are using most um but there's there's lots of uh there's a few private Alternatives and there's a bunch of Open Source Alternatives I think a lot of um I think again it depends on the use case but a lot of use cases we see are some type of kind of like retrieval augmented generation or grounded generation and so that generally involves a vector store that generally involves an embedding model um and I think another interesting thing is also uh uh and yeah and there's a bunch of there's a bunch of different things that you can choose between there um but um I think another component is actually the UI for it as well um and I think like uh there as kind of mentioned like um alums aren't great for every use case and a way that you can make it better besides improving their line model is to improve like the UI and the ux of the things as well and so I think things like showing intermediate steps showing the retrieved documents help build some of that trust help people who are looking at answers validate whether it is actually correct or not and so I think um yeah there's there's a bunch of uh there's a few kind of like existing players like radio and streamlit that have kind of like got into the space with some really interesting templates there's also new players like chainlet which offer kind of like this intermediate thing I think the UI is maybe like the UN the most underexplored and underdeveloped part of the stack like I feel like everything else is like relatively kind of like standard but I think like everyone's doing chat and chat's hopefully not the only UI for these things and so I think yeah I think there'll be some interesting kind of like UI Frameworks to emerge to help with a lot of this awesome anybody else want to take a stab at that yeah so you know when you think of um you know some of the co-pilot style um Integrations right so a lot of organizations that we work with um immediately from a you know Frontline worker perspective um in financial services or uh you know my background I did 10 years at the National Security Agency we work with tons of analysts and what they talk about is even if you develop a new application that's outside of their workflow even if they have to do that concept switching from one application to another they'll never use it right because it's out of their workflow and when they want to make a decision on the data and it's out of context they won't use it so one of the things that is really important for these analysts that may use one or two tools on a day-to-day basis and they don't use anything else is really trying to figure out how do you integrate these large language Model results directly inside of the workflows of the one or two tools that they use on a day-to-day basis so um to your point I think it's really important not only exploring kind of the one-off applications for the conversational AI but trying to integrate the results directly into workflows of those users I think there's also like adding on to that I think there's another type of ux which is interesting to explore which is maybe like the no ux Paradigm or basically just run things as background processes and I think that's also relevant for this discussion around risk because if you think about a chat there's a very real like expectation on the latency that you can induce but if it runs in a background process like you know there's less of an SLA on like an email or something so you can maybe do a lot of fact checking with like there you know there's a bunch of research using llms to fact check kind of like themselves doing some reflection steps and so for ux's where it's just in the background it's not in the foreground you can maybe do more of these and increase kind of like the reliability with that as well awesome all right so uh I like it because normally I'm doing this virtually and I can't tell if somebody wants to talk or not and so then I got to call them out and then somebody's like oh no or and here I get to see the raises the microphone to their face and in their head and so then I will say this okay so so here we go I'm sorry um I think I think there's um I think we just have to be mindful of llms like checking themselves right I think it's um tell us more I mean I mean I I I think that we um I think that you know we're aware of the fact that like you know there's they're resisting using llms and when we're using LMS to uh you know when we're sort of using out of the box LMS to test um you know to Tesla levels and I think that um you know we've seen evidence of of cases where where that um where that can go where that can go wrong so um so I think those that that is yet yet another risk you know that we have to be uh mindful of um so so basically don't rely too much on the llm that lied once May lie again yeah I think we just need to be yeah we need to be thoughtful about it we need to build like kind of um tools that are purposefully built for uh for these tasks and so that brings up the next point which is kind of like common pitfalls that you'll see and I think that one is a clear case of people saying okay well if we can just get the llm to evaluate its output then we should be good and later down the line you realize that you're not good what are some other things that you've seen along the way and you're like yeah this is being taken as good practice but may not be that useful and I see a few of you kind of having to think this wasn't on the pre-recorded questions so now they're starting to go like hey he's throwing a curveball um Ben you think of anything I I think one one big area for a lot of organizations is there they're they're really trying to increase speed to market right now for these large language model capabilities and um you know we were um uh you know in a uh a meeting with a CEO and um he was talking about you know the importance for him to accelerate the AI machine learning and uh he talked about his board of directors and they told him that he needed to take part of his pay and figure out how to increase the amount of AI that was in the product for his organization right so so there's a lot of emphasis on trying to speed to Market and what we've seen uh in two different types of organizations and highly regulated organizations and financial services and top secret classified organizations in the intelligence Community is they do a really good job of running observational studies first right so even if they're trying to increase the speed to Market they're going to slow down to do an observational study first where they're Gathering feedback from a handful of of users inside of the organization where they're tuning and retraining before they deploy to a a larger you know hundreds of thousands of end users that need to use a technology apology so even though that these organizations are trying to you know accelerate how quickly they're getting the AI and machine learning results from the larger language models in the hands of those end users is still extremely important to have that human machine teaming aspect up front where you're starting to gather that feedback because ultimately if you can't kind of gain that trust from that study and then scale it in a way where you have that trust across the organization you won't get adoption anyway so it's really important to kind of have that two-step process excellent so I want to be mindful because we have just a bit of time left and I want to make sure that you all can ask questions so I see a few hands up already I'll give my microphone over it's going to be like the equivalent of putting myself on mute and I'll let you all ask the questions now okay so as you're doing this calculus the risk versus reward of the of the large language models the calculus risk versus reward along the way with the training at what point or what are the top two or three either way that makes the clear to you that the risk is exceeding the reward and you need to take a different path or that the reward is exceeding the risk what's the the top one or two or three in each direction that leads you to go one way or the other that the risk exceeds a reward or the reward is exceeding the risk I mean I think the honestly a lot of what we see is that um there's not great like I think people are struggling to find product Market Fit For A lot of these applications in general and so I just don't think the reward for a lot of things is that high like especially if we're talking about some of like the tool use and agent stuff that can introduce some of this risk like there are a lot of um there are a lot of there are a lot of demos on on Twitter of using agents to do various things but I just don't think they're good enough to be put in production and it's not necessarily even because of the risk at this point I don't think I think there's just a gap in going from prototype to production and so um yeah this is like a roundabout way of saying I just don't think the reward in general is that high for a lot of these more complex applications which is generally where where link chain kind of plays in so that's where I have more of the visibility yeah and I think it's not different from what we you know know from like basic ml right if you can do it with the decision tree then do it with the decision tree you know if you can do it with linear with linear regression great do it with linear regression don't do it with a you know um a deep uh deep learning model um and I think it's not different here so I think the the first question that people need to ask themselves is like you know when they're when they're using an application do I really need you know the state of the art llm with you know all the you know risks associated with it maybe I do maybe I don't um but um but then you know basically try to find like I think the simplest model uh that you could use and this is where I think the open source Community uh is going to be so powerful right because I feel like so many people are you know now are very empowered to build their own llms for their very targeted specific applications and I think that's gonna um ultimately what what that would do is uh reduce risk I think time or one more question thanks uh first dimitrios you're more fun than two weeks ago at ml Ops so amazing um so I'm from the hedge fund community and as Benjamin pointed out there's a lot of regulatory like it's just very strict right um and there's been a there's balance there's a pressure from investors on like are you using gbt yet um on the other hand uh SEC and you know other managed directors are like maybe lean on the earth from the risk side but uh I'm thinking you know there's a lot on top risk like untapped opportunities like vector search or summarization there's a lot of fundamental research you could do or a Quant so I'm thinking like traditionally ml you do a lot of confidence intervals right that or credible intervals that is more important than the actual point estimate or the prediction so I'm thinking here what are your thoughts on some kind of probability Beyond just like the token next word prediction but just the actual here's the answer here's the closest match why is that actual number some kind of probability that is intuitive for the Layman to kind of understand uh so there's some like interesting stuff around asking like the language model to Output kind of like probabilities but I think most of that has shown that they're not like super well calibrated I think that's what I've seen like empirically as well um I think they're just like overconfident in a lot of cases and they're also biased towards certain um like around how this gets back to the point of having alums create themselves but they're biased towards like certain places in in the thing um I mean I think like what we basically see people doing and this is a little bit this is kind of related to getting like a confidence estimate is like when when you start you start from scratch for most cases I think that's the power of the generative models they can do everything like yes you know you could train a classifier but instead you could just like prompt the language model in this way and you can get up to speed really fast the issues then is like how do you know how well it's doing and I think what we see people doing there is just building up a data set of like test cases over time and then running against that and so when you change the prompt or whatever you kind of just like look to see what happens on your 10 15 20 different test cases and I think that's it's it's not exactly what you ask but I think like that's how we see people getting a general sense of how it's doing on a task in general yeah yeah a combination of human in the loop combination of like uh testing basically um and this also gets to like uh like I know we've talked about kind of like L alums grading each other as well and I think it's not perfect but I think one of the things that it does do is it can be like a Guiding Light to for where the human should focus their effort so if you have kind of like a thousand data points maybe you don't trust the llm I think I don't think anyone who's doing a evals trusted completely to give a grade of like 0.93 or something like that but I think what it can do is generally highlight data points and then you have a human go in and look at maybe like 10 data points instead of like 100 and I think we've yeah I mean I I think like we're still so early on in in this journey and I think a big part of it is just getting that iteration speed up and so having a way to measure that confidence is really important in the way we see people doing that is just having this data set of like 10 20 examples just for um you know we have um uh with there's actually a lot of good academic research on conformal predictions that do exactly this that give you not only prediction but give you like the confidence around that prediction and there's there's a very solid theory behind it um to um you know I think that um we haven't seen it as much applied uh for you know large language models but I'd be surprised if we you know in the you know in the coming you know machine learning conferences we wouldn't see uh you know that line of research got adopted to uh to this so so I think if we just sort of continue on seeing what what comes out of the academic Community I think we're going to see exactly that and I'm sure that it's going to be very interesting all right I guess that is it everybody let's hear a huge round of applause for our panelists thank you all so much for joining us tonight