LLM Avalanche: Panel: Enteprise & Deployment
Recording: LLM Avalanche: Panel: Enteprise & Deployment
foreign this is pretty awesome because normally I do this kind of stuff online and getting to see your faces is oh it's a new experience for me so in case all of a sudden I'm staring like a deer in the headlights you know why no we've got an awesome panel and I'm very excited to talk with these legends that are on stage with us you can see the wonderful panelists that we've got I'm going to ask them to give a quick run through of who they are because they have long list of credits behind their names and I couldn't remember all of it so I'm going to go ahead and ask Inez to start us off and let us know who you are and what you're doing and then we'll get into this panel and we will leave 10 minutes at the end of the panel in case anyone has questions so you know think of questions all right there we go and that is kick us off um hey everyone I'm Inez I'm a co-founder at member station which is a startup that does data transformation data preparation with Foundation model technology um and before number station I was doing a PhD at Stanford uh in Ai and specifically in using AI to automate some data work like Knowledge Graph construction towards the end of my PhD we realized that Foundation models basically could replace all that work that we've done but that was also exciting and that's why we decided to start a number station and super excited to be here today great so I'm gonna zigzag through Jonathan can you give us who you are and I think your life's changed a little bit in the last couple days so I'm Jonathan Frankel I'm Chief scientist at a startup called Mosaic ml that will no longer be a startup in another couple weeks um all of the chief neural network scientist at databrick shortly which is very exciting I train neural networks and my focus is on doing so as efficiently as possible I spend my PhD at MIT studying that and now at mosaic and soon to be I guess a data bricks we're going to drive down the cost so everybody can train a foundation model for themselves there you go all right Emmanuel go ahead yeah let's give it up for him and huge congrats on the acquisition that's that's awesome Emanuel um hard to top that in terms of a big day so I'm impressed you're still alive I was I was in machine learning for a while most recently I was at right before I led some efforts on the fraud team and I'm part of the technical staff at anthropic where we work uh you know on excellent and Andrew can you give us who you are yeah I'm Andrew Dai I'm a principal software engineer at Google deepmind and I've been working on language models since around 2015 when we were building uh small models with these uh ancient things called lstms and back then we would like create chat Bots and things like that but we'd never imagined that in less than 10 years everyone is using them so these days I collect the Palm 2 project and now I'm working on Google sge the search generative experience and also Gemini excellent and sudeep you are going to finish this off all right uh hi I'm sudeep I'm the director of engineering at cohere cohere is a foundation model provider um which is based out of Toronto but we do have Offices here as well as in London uh before kohir I used to work for Google as a researcher and Google brain as well worked on ML Pathways which was then uh used for training large language models like um Palm at cohere I lead a couple of teams one of my teams is responsible for uh inference so deploying our models in whatever environment we we wish to deploy them in as well as uh for fine tuning so customizing our language models for uh especially Enterprise customers excellent so now you understand why I said they got to introduce themselves now one thing that I want to start us off with and this is a panel on Enterprise and deployment right so Enterprise I think there is a very big factor that goes into using machine learning and AI with Enterprises and that is what is the ROI and how much money does it cost everything around money we want to get very clear on so Jonathan I'm going to kick this off with you and I know that you mentioned something around the brass tax of how much these things actually cost and how we can look at cost and so uh when it comes down to the costs I know there are huge questions whether that is how much it costs to actually train a model or fine-tune a model or whether it is how much it costs for the resources that I need to actually be serving these models or paying the engineers so with all of that in mind how can you have like a framework or something to go off of when it comes to looking at the costs I think it's actually pretty simple at least for a lot of the folks you're seeing here we've already built the models for the most part or at least built the infrastructure to build the models and that cost is baked in So at the end of the day everybody up here should be able to give you an exact number for how much it would cost to fine tune and deploy their model I can give you some exact numbers if you were to take our model and you want to Define tune it per billion tokens you're talking about seven hundred dollars that's not a lot um if you wanted to you know one engineer a month let's say call that you know I don't know twenty thousand dollars maybe plus or minus so if you're willing to invest one engineer a month worth of you know worth of investment in fine-tuning one of these models to customize it cool you're getting you know tens of billions of tokens worth of training that's really significant on your data I'm curious actually to ask my peers here how much would it cost to customize the cohere anthropic models I I don't know genuinely I'm thank you okay so um I'm not going to answer that exactly I tell you why for the acquisition I'll tell you why though um for one um I think there are a lot of considerations that go into determining what the eventual cost is right for instance uh what kind of fine tuning you want to do how much of the parameters you of the model you want to update but then what size of the model what kind of Hardware you need to run parameters all the parameters of the model right so I I don't know that off the top of my head because those are not our model sizes right so um uh but uh I think when we talk about rois I think it's important to actually um consider the return side of the equation as well right so um let's say that you actually deploy a model in a healthcare setting which enables a doctor to write up a referral really easily so a task that used to take him 30 minutes now takes him two minutes because he doesn't need to draft that letter he just needs to verify what the language model produced similarly a lawyer who may be charging like 500 to a thousand dollars an hour can probably uh you know be much more efficient at his tasks so uh it doesn't really matter if you spend like a few thousand dollars or a few tens of thousands dollars it on like you know fine-tuning a model because that cost easily gets um uh overcome by the uh much significantly larger returns that you can actually get in terms of human hours saved through these language models for instance excellent all right so but I I still have to ask is it a few tens of thousands of dollars then as I mentioned I think uh no it should actually be a few hundred dollars for uh some of our models awesome okay that's brass tax so all right now we we've got a few more questions cruising along this cost idea and Emmanuel I'll kick it over to you because I know that claude's got a gigantic context window and one big question that I always think about is when do you look at using that whole context window versus not using it and why would you use it and especially knowing that each token is going to cost you a little bit more [Applause] I like that we got this far without me realizing okay um I train models I don't do microphones uh but anyways um what was I saying right 100 000 tokens is the limit for Claude which is roughly the length of like an average book um so when do you use that versus when you don't I think there's like a question of like for most for many things that you do you can rely on like you know internal knowledge and just like ask the model to do a thing and then you don't really need like a very large context window but there's like a lot of tasks where you need to like pull in knowledge and that's kind of like tasks where you've often seen things like Lang chain be used for example right where like you're like okay uh I have um let's say like I have all the data bricks docs and you know I want people to answer to be able to like answer questions about these docs um instead of like retrieving these docs and like searching for them which usually like like takes some time um and can be error prone you could like put the whole docs uh in the context window and like ask a question about it um I think that generally like long very long range uh chats is like another exciting thing for me personally where it's like you could chat with like a chatbot for years and have it remember like everything you've ever said um and the other thing I think that we like call down to the demo which like blew my mind was like you could just paste like the financials of like a company that just announced they're like you know like uh quarterly results and in like five seconds get a nice summary of like what happened so anything that's like long document summarization that sort of stuff is a good fit awesome and Inez I feel like you also had something along these lines and the techniques and all that good stuff yeah it was essentially about how do we inject knowledge in these models and for us we work with structured data you mentioned documentation of tables Etc and how do we make the models aware especially each organization has their own terminology for financial terms Insurance terms and the public models don't have that so these context windows with retrievers is one way to do it so I like to to call those inference time enrichment because the model weights are not changed but we can also do the fine tuning piece as well so when there's too much and we can't fit everything in the context the models are really good at synthesizing information and and summarizing that uh through the weights so there's research on both ends like one making longer and longer context windows and one on fine tuning on domain-specific data so yeah excellent Andrew I saw you Bob in your head you got anything you want to add to that um yeah I want to say that like long context Windows to offer more applications so especially in things like coding you might want the model to know what's in your header files rather than just like image guessing what's in there that's usually useful so it basically enables a wider range of applications from like long book summarization to financial report summarization as some of my other panelists have mentioned okay so this is Enterprise and deployment I want to do a little bit of the deployment part and then we may get back to the Enterprise but I'll tie up this and this whole Roi and cost uh piece because there is something that I just did this whole survey with a lot of people in the mlops community which I run and a lot of people were talking about how when you're measuring Roi it's it despite despite what Jonathan says it is still a little tricky to figure it out and whether that is the time that an engineer is putting into having to deploy it or the resources that you need the GPU costs all that stuff comes into play and then you look at how are you where are you taking that from so is it like cost of goods and how does that work somebody's phone's blown up back there that I can tell yeah take that call so we are um when it when it comes to the ROI it's like where is it coming from just to clarify your question are you looking for the r or the I like is this trying to calculate where the I comes from like calculating how much it cost to actually deploy this or is is that what you're getting at no I think what is just not clear still for people um besides you is that the the whole Roi is is how much how many resources do I actually throw at this and how can I figure out how much money I will make off of this tool or this feature if I add AI to my my product can I charge more money because I have this Ai and if not then I am still selling at the same price but I am paying for the AI does that make sense yeah I think you know I'm guessing we would all have the same answer to this which is you start small and work your way up the only way to really find out is to try and very tiny Investments and seeing what kind of return you get allow you to justify making bigger Investments running small experiments a b testing it's honestly pretty similar to any other change that you'd make to a major product you need to invest small work your way up see where you're getting signal and continue improving from there I'm curious if other folks have had a different experience um yeah I was just gonna say like it's it's funny because uh um I was just talking about this uh conference earlier that you were running um but like so I used to run like an ml education uh program and so yeah like tips like these were ones that I would share often and I think one of the things that like I drilled into people's brains and I'm sorry for that is like you know just always start with like linear regression or like get some like kind of like simple like pre-trained um you know bird or whatever and I feel like one thing that I observed myself recently is that now you have these like very very powerful uh General models uh and that's kind of like never really been the case before in the sense that they've never been generally enough that there would be a chance that they would work for your like very specific use case but now you know you can use like you know cloud like a cohere model like tragic piece I'm open source model and just like kind of like get a very quick estimate of like How likely is this to work for your use case now of course like that model won't necessarily work as well as like something that you'd like train yourself or fine-tune yourself but I feel like that piece is something that I've made you know a recommendation uh update on when I talked to like founder friends or like people that are at tech companies is provided that you know there's no kind of like uh data privacy concern to do that immediately like try one of the large models and like give it the instructions for the thing and see if like that works and if that works you have like a really good um at least I estimate right you're like well this is how much is going to cost me per call to do this thing is this worth it for me or not and then you can discuss kind of take to The Next Step but having that first step that takes you know 30 seconds is something that I had to update my prior on personally you got something um no I pretty much like agree with whatever else was said here um yeah I think the first step should always be just maybe like use a SAS API right you pay per call it's easy to integrate easy to experiment with um I think uh on the other side again on the return side it really depends um it's it's not always as easy to measure the impact yes you want to do a b x a b experiments and a b testing as uh to evaluate it but assigning a dollar value on the return is not necessarily always as easy right for certain use cases of language models let's say uh for productivity um uh increases use cases right if you you're if you're using copilot if a marketing person is using it to draft a copywriting email then uh you may be able to like measure it in uh employee hours saved for instance there are other use cases for instance if you're using it for Content moderation of your gaming chat room uh for instance then it it's it's harder to exactly quantify right there I guess the the quantifiable metric would be user happiness reports whatever you are monitoring whatever you're using to measure user satisfaction and then you need to assign like a dollar value to that so uh it's not trivial but yeah there are ways and proxies for determining it sorry and I just want to say that for some things like a customer facing features Roi might not be the right way to think about it so overall like this technology might reduce your margins but if you don't deploy it and put it in front of customers other companies will so uh my in the end be a question of survival awesome and there are other times we're just hitting an API doesn't work and Inez I know you are in that camp where you can't just hit some external API can you explain how you look at that and how you think about taking it one level deeper yeah and and I agree with uh what manual said on the first step I always advise as well to use apis and prototype before actually building something um for us in the data space um and data management there's a lot of uh personalization that's needed to understand the domain-specific data so Finance Insurance all these domains it's really important to do that fine tuning um as well as privacy is a big issue a lot of the organizations are not okay sending their data to a third party even if sometimes the third party company is not using the data for training they're still not comfortable with that so having a way to personalize the models for specific Industries when they have very sensitive data is is really important only in certain verticals of course like if it's e-commerce sometimes it's not as sensitive as Healthcare but that personalization and distillation of model is is key excellent so I told you we were going to go with deployment and then I went on a huge tangent there so we're going back to deployment now and Studio I'm going to ask you the different factors that you look at for choosing the right deployment strategy and going back to that um that survey that we did I know a lot of people talked about how the current infrastructure makes it very difficult to serve these gigantic models and so how do you look at that when you are trying to deploy right so I I think when I uh suggested that question I guess it was not necessarily more from the point of view of how we deploy but how a customer would choose different deployment options uh when they are considering potentially different providers and different deployment options that are available from different providers so usually the initial first thing that anyone would try is a SAS API right it's easy to use and you would not even consider customization right the first product feature launch that gives us an llm would probably just use one of the SAS API you would potentially choose a product or a feature that doesn't have stringent privacy or security restrictions if if you're a product actually allows for that um after that once you're happy then I guess the next step is doing considering whether you actually have proprietary data that you can then leverage to improve the performance of the models and uh there you need to consider different strategies like whether you want to use a retrieval augmented generation model so that you can retrieve those relevant context from whatever external knowledge sources that you have um or you want to just bake in some of the knowledge within the model itself and if you want to bake in some of the knowledge then you then consider fine-tuning so there's basically a spectrum of ease of use and using things out of the box to things that require significant investment and tuning to get the last bit of performance and cost savings at the other end of the spectrum so yeah when Enterprises should like consider what's the best fit I'm curious and this is this is an innocent question for the folks who deploy big models is it more expensive if I were to customize your model to deploy it like what do my deployment costs look like if I use the generic API versus if I wanted to customize the model and then use that does that change things I mean I I don't see a reason why it would um but but I think like uh maybe just to go back on like the actual question like I think that wouldn't change really um like one thing that that I think would like make a difference that I was thinking about is that like I think using one of those apis like makes a lot more sense also if you don't have existing infrastructure in place so like one thing that before taking off my anthropic hat before I was in a topic I was at stripe uh and at stripe like we had really good ml infrastructure and so like adding one more model you know like was just the cost of figuring out how to provision that machine and like we already had like kind of like same you know feature storage feature deployment like model provisioning model monitoring et cetera Etc yada yada um if like a lot of companies uh you know the like llm based application is like your first mm application then I feel like uh building that whole suite for it is a heavy lift and so that's also like one thing where you should consider in your deployment story like hey can we reuse a bunch of things that we already have and in which case like yeah like add another model to your like Suite of models and then judge them based on performance versus like do I need to build this whole entire stack that there are entire like field studying like the ml Ops field anew oh all right uh so I think that we can keep moving when it comes to like this idea of Open Source first closed I get the feeling somebody on the uh panel is going to have a lot to say about this and so um Jonathan I want to kick this over to you and just think about how you look at whether or not it should be on or how you think about open source first closed and then if you advise hmm closed ever or uh or open and when and why I you know I I suggested this question because I think the open source versus closed Source question is when I get a lot from customers I think it's completely the wrong question um but I hear it a lot and it's worth kind of addressing and I'm curious to hear what other folks think for me it's a really transparency and control versus lack of transparency and lack of control there are open source models where we don't know what they were trained on and you have about as much transparency into that model as you do into Claude or coherence models or you know gpt4 or what have you um there are plenty of Open Source models where you do know everything about that model and you probably don't want to touch it because you do know everything about that model and that should make you very nervous um with the closed Source models you really legitimately have no idea then there's a spectrum of control versus lack of control with an open source model you do have an enormous amount of control over what you do with it with a closed Source model you may have no control you may have a limited degree of control and so I think the real question for anyone you know again I love the open source versus closed Source question because I think it's the wrong question but people really mean transparency versus lack of transparency and control versus lack of control and it depends on the application if the model is good oftentimes you don't care what's in it if you're in a medical application you may actually care that the model is trained on WebMD or what have you you know that can matter quite a bit and so to me those are the two right questions there's no right answer but you know as an application designer you you probably have a good sense for whether you need control or not or at least you can try and find out and you probably have a good sense whether you need to know what was in this model for very specific reasons or whether you really don't care as long as it's good sure yeah I must say I think that was like a surprising reasonable take that's good thank you um now like seriously because yeah I think I agree with at least like well I agree that it's a bad question I also agree that like with with like kind of the transformation transparency taken like uh control I think if that uh matters to you I think like the other thing is this to me feels a lot there's a lot of Holy Wars uh this to me feels a lot like the uh you know if you're old enough uh Cloud versus like your own data center uh holy war was a thing for years where people were like well like you'd be a fool to like run any business on top of like a data center that you don't control like you need your own racks and your own machines and like in some cases I think that's still true today uh right like for some applications you still need your own racks and your own machines in many cases you're like yeah like as long as the machines run my code and I'm happy with it like that's good I'm not saying that the limitation will be the same for for models I actually think it'll be quite different but I think that there's there will be like both for a while and there'll be reasons to use either the last thing I'll say um and I think it's like obviously now putting my anthropic back on like one thing that we care a lot about at anthropic is like safety and so if you're like generally concerned with kind of like giving access to like any model to anyone and imagine like maybe like not today's model but like a much much much more powerful model that's an argument for also like in some cases saying like Okay for some capabilities we'll like limit access to those uh but then I think that like there's many models that don't fall within that uh distinction but I want to call it out yeah okay um I pretty much like agree with most of what was said and that was actually a fairly reasonable take right why do you sound so surprised because we heard the prior ones um so I do believe that you know it open source and closed Source uh models will continue to exist uh in the future um you know cohere itself has a cohere for AI which is a non-profit research arm which trains models uh based on community contributions uh the data that is used to train these models is open source the models themselves are open sourced as well but we do want to also have a business and we do want to solve specifically for problems that arise in large Enterprises whether it's on-prem deployment whether it's actually developing a product that actually solves real use cases uh instead of just focusing on the model itself so for all of those cases for all of those reasons uh like closed proprietary models will continue to exist in the foreseeable future um but yeah I all for open source models and a thriving open source Community as well yeah I I agree with uh um what's being said before and um there are two Emmanuel's point I do think like the cloud versus like your own server data center I think it's one way to look at it because uh right now these models are static right you train them once um they're fixed and then they're deployed which fits very well with an open source model but that doesn't mean uh this will be the way in the future right like you can imagine models might be continually improving or continually updated and in those cases probably the closed Source models will be more up to date than the open source ones which which require you know a lot of investment to to keep refreshing yeah and also aligned with everything that was said I think there's also the aspect of size um the very very big model can only be trained by big organizations like anthropic openai et cetera so I think we're going to continue seeing the very good AGI models uh held by those organizations but then there's also that wave of Open Source models that are much smaller Recent research showing that one billion with very good data can get me very high quality so I think we're gonna see those two trends at least that's my hypothesis very big models solving AGI general purpose held by those very big organizations and then a wave of Open Source that are going to be distributed and fine-tuned by each organizations individually awesome all right so now we've come to the questions so hopefully after this thoughtful discussion you all have some questions on your mind go ahead and raise your hand if you have any yeah I see some in the back I'm just gonna ask you to stand up and shout your question at us because oh no we have a microphone coming to you all right perfect that's even better then you don't have to yell at us uh that works all right so we've got the first question coming right up yeah go ahead and grab that um so I think my first question is more around what are the right approaches to kind of like fine-tune your model right we heard a couple of approaches here one is like should we fine-tune our models should we like you know use techniques like RL HF should we prompt our models to kind of like make them more conditional to the context right even there I think there are questions whether we do retrieval augmented generation or we like enhance the context window to a point where you know everything can fit into the context window itself so is there like a decision tree or is there like a way to think about it where Enterprises can make these decisions more easily yeah I so in my opinion uh start with the simplest and the easiest um and then uh fine tuning is probably actually at one extreme um and even before fine tuning uh I think one of the important steps that people usually gloss over is uh how do you actually collect the data that you want to fine-tune on and how do you ensure that that is actually high quality data um and um uh that bit is often challenging like to basically get high quality data that has been human reviewed um uh and get sufficient volume of it to actually be able to fine-tune a model uh is often more challenging than just you know running the fine tune itself um so in order to avoid that I think you can get a lot of performance with just retrieval augmented Generations right um because uh that addresses a lot of the problems that usually happen with hallucinations um it with data freshness because you know you're retrieving the documents from an external memory which is always guaranteed to be fresh you none of knowledge is baked into the model itself so you need to take into consideration all of those factors instead of just going for fine tuning which is which may be very applicable in some some situations I think another way to think about it is the size of your fine tuning set so if you were creating your own set in an afternoon then maybe you have some like 10 examples and that's good for a few short learning but if you have a bit more maybe that's better for prompt tuning and probably if you have like on the order of thousands then you want to do something like fine tuning the answer to most of these questions of what do I do is cheapest to most expensive and easiest artist this is all in empirical science and we don't know the answers things inexplicably work or don't work and it's really hard to predict in advance a lot of the time so do its cheapest and work your way up the ladder um maybe the last thing I'll say is like people are surprisingly bad at prompting is what I found so it's it's worth like reading there's a lot of good open source docs on like I think like yeah like whatever open AI the provider's anthropic maybe cohere have one but there's also like GitHub repos with like prompt tips and it's like there's a lot of stuff you can do that you know like you might think a model just does does not do your task at all and then once you apply like 10 to 12 uh tricks it just does it and so I think that's worth doing before you do a bunch of work yeah and there's also piggybacking on that there's also a lot of good open source Frameworks out there that can help you like something like promptomize and then you can test your prompts and see which ones work with what models so we've got another question back there yeah hit it yeah I'll I'll be curious about what's your customers take on privacy and maybe even the Federated Learning Systems about that so maybe you can speak about that I mean I think those are two separate questions really um when it comes to privacy there are certain industries where you simply can't call out to an API the model has to be within your control and you need to know all the data that's going in and out um in a lot of those contexts there's no substitute the model has to be somewhere where you can control it and where you know the data is not going anywhere else or being stored in those cases there's nothing you can do Federated learning you know when it comes to there are a lot of different approaches that have been proposed for trying to handle privacy in various cases federal learning is a whole separate example where you know you really need to separate into training and inference Federated learning is a training technique where if you really need to combine a lot of different data and perhaps you don't want to see what anybody else is providing data wise but you do want to be able to see updates to the model we don't really have that strong of guarantees around things like Federated learning even differential privacy which hypothetically gives you a guarantee it's really hard to concretize that and make it explicit you know you have a certain amount of noise but the threshold between did I add so much noise that I ruined my model versus did I add a too little noise and the data can be extracted from the model it's all in empirical science and to me it's still a pretty dangerous game to play you don't really know what you're getting at the end of the day your best guarantee is if you control the model you train the model on whatever data you want and nobody else can see that model except you that is your only way to get a true guarantee right now yeah I was gonna say I feel like Federated learning is like fine-tuning but worse right now and it's it's value prop it's like if if you're thinking about it you should probably just fine tune um for for what the first one you said yeah agreed I think there's starting to be like uh some models that people like deploy within their like clouds I think opening on something about like Azure and whatever I imagine there'll be more of that where it's like you have some like model and it's like deployed in your Cloud um or you have your own model yeah I think there's basically a spectrum of choices right so there's the SAS API which logs potentially some of your data some of the providers will let you opt out of data logging but it's still a SAS API then we have Solutions like AWS sagemaker which are hosted platforms and for instance like cohere models are available there and you can spin it up and still get benefits of a managed service and it's completely private no data leaves uh it leaves your your environment on top of that there's like virtual private clouds which is deploy our language models within your virtual private clouds and on the extreme and it's fully on-prem Solutions which are basically air gapped environments these are mostly like financial institutions or really highly sensitive Enterprises who do not want any connection to the internet and we have Solutions in that space as well so it really depends and obviously like the complexity of the deployment kind of increases as you go along from one end dot to the other end but it really depends on which which of these options is the right fit for you um so hang on sorry sorry so my friend invited me to here so I don't really know what's going on I don't I barely use the internet but what do you guys see uh like AI or like the world in 10 years from now all right I think that if you would have asked most of the people on the stage what they would have said five years ago about what the world would look like now I looked at gbt2 and I said oh that's cute um that's a very silly thing why did they train something so big then a couple years later I looked at gvt3 and said oh that's cute why did they train something so big um and here we are predicting the future is futile um and the only thing I can guarantee is you're going to be wrong and the most extreme predictions are unlikely to be true that's it yeah uh I said yeah like I said that the model after GT2 would be a huge waste of money would never work and I also said the iPad would never work so like don't ask me anything so meanwhile how are all the crypto folks who predicted the future doing so you never know it shows how quality this panel is debate us on that question I guess uh did we have more questions yeah I see you see one person here yeah looking at you yeah and Inez I realized we did basically the uh in person version of a mute button on a virtual panel where you took your microphone so we'll let you have this one hopefully it is so frankly I just follow up the Federal learning questions so it seems like you're essentially pretty down on the federal learning part but there are many cases that you just cannot put data together for example Hospital have different if you want to train something with several hospitals in their privacy data the patient's data will not move and if you're in the same company in the different Europe in the American China or Asia the data cannot be moved out of the region so with all these laws privacy laws so you have to use a federal learning to train these models so I it's kind of a surprise to see that you're pretty the model is the data the data is in the model if you put the if you put the data together in a model you have all that data in one place within the model and that data can be extracted from the model you cannot divorce the data from the model and so if you do Federated learning you have brought all the data together it's just in a different form got it thank you hi thanks I was curious um what different approaches you have for evaluation Frameworks and if both quantitative and qualitative and whether you know how how automated is that um seems like a really interesting topic and probably an area where you need to develop a lot of tools I can I can speak for data transformation applications so it's very much human in the loop in the prototyping phase so user validate outputs and we can like monitor that interaction and understand how well the model is doing based on whether they clicked something or not so it's very interaction based and in production it's it boils down to classical mL of I have a ground truth set and I'm measuring quality on that ground through set and making sure that the model performance is not degrading over time Etc so more like classical ml Solutions once the model is deployed yeah I think evaluation of large language models is like a fairly complex and hard problem which I don't think has been solved yet there's obviously a lot of academic benchmarks that you can use to measure the performance and do relative comparisons Stanford runs the helm benchmarks which is again like a basically a neutral evaluation of a bunch of large language models that are provided you know that were evaluated I guess across 60 different task verticals and have multiple providers listed there so you can get a relative evaluation of how each of these models is performing on each of these tasks but what we have seen is that there are two things that we usually Focus much or give much more importance to one is um specifically when we work with Enterprises we want the models to solve their specific tasks so how our model actually performs for those specific Enterprises so we develop custom evaluation Suites for each of them and then the second is human evaluations we have internal Army of evaluators who evaluate our models for various various aspects and we have seen that that human evaluation is really the most indicative of how well the model is perceived true but here's another thing I never thought I'd say uh which is that now uh we're getting into the the weird bizarre world of you can actually ask your model uh to evaluate the model and that sometimes works better than humans um initially when I thought I saw these results I also was like oh this is stupid it will never work um but it turns out it makes sense especially for like the task you're evaluating is something that would take a human like let's say like longer than a quarter of a second like you'd have to like read something and be like oh well you know is this actually like an effective summary or whatever it turns out that like human Raiders are human and they get tired and they make mistakes when they write like you know 100 things in a row and so not for every case but for many cases you can like actually um you find that models especially large language models are better at grading than they are generating so the same model that like messed up half the time on the question you asked it can get like 95 you know on like actually giving you the golden label uh on every question and so in many cases that's something that's worth trying and also makes no sense yes so I would I would agree that uh that that is uh I would say that's like a more extreme take on the on this um I think like if you have a model that upgrades itself and it says it's 100 then you know that doesn't say a whole lot but if it says a great greatest self and it gets only 50 that definitely says a lot right that means your model isn't great or isn't great at lying to itself um but so you can use another model to grade your model doesn't have to be the same one right uh but but but that's uh also there's another point there which is like most of these models are trained on very similar data right they all trained on web scripts and it's very likely that the deficiency in one model will show up in another model um so I'd caution against that the other thing I would say is like uh be careful when you look at academic benchmarks a lot of them right now are tainted um so there is like quite a significant degree of contamination in all these and I think like the people should think about like doing things like rolling benchmarks or especially human evals are a bit more resistant to these kinds of things but just take all the eval results you see with a grain of salt yeah I've I've seen the throat cutting gesture which means I think it's time for us to wrap up um but I mean the I think they're really a couple lessons worth taking away I think I've agreed with what everybody said number one is there are plenty of benchmarks out there um the academic benchmarks are sometimes vaguely indicative of when a model is garbage versus not garbage but that's about it often these benchmarks are reasonable at like you know this model is trash and this model is not but aren't going to give you fine-grained distinctions between models that are various degrees of good there are a bunch of interesting leaderboards out there and they're very interesting but I don't know if they have much more information than that but I think the you know to maybe come back to the root of the question if you were going to do any large language model related tasks or any really deep learning task today evaluation is where it all starts until you have some measure of evaluation you're not even doing machine learning you may as well not bother the thing that scares me most is if a customer came to me and spent 10 million dollars training a model and then said afterwards so is it good that is the that is the nightmare scenario you know when I wake up in a cold sweat it's because I had a dream about that happening um but in all seriousness the evaluation is where it begins how would you evaluate a group of humans who are performing that task if you don't have anything to go on there you need to begin you know to figure that out before you should even try to train a model or try to use it for anything that's it [Applause]