Devreal

SBTB 2023: Nate Slater, What problems should enterprises actually trust an LLM to solve?

SBTB 2023: Nate Slater, What problems should enterprises actually trust an LLM to solve?

Recording: SBTB 2023: Nate Slater, What problems should enterprises actually trust an LLM to solve?

I've given a lot of talks over the years and this is probably like the nicest uh Ballroom uh that uh that I've uh had the chance to speak in so uh before I begin I just want to tell you a little bit about who I am and why I'm here today to talk to you about uh um generative Ai and uh how it can be effectively used in the Enterprise so uh my name is Nate Slater I'm one of the founding team members at flip AI we are uh a devops llm company that's uh basically training llms that are capable of um interrogating observability data and producing a root cause oops we just lost hold on one second here we back okay um so uh we we have a product that's actually able to use llms to produce root causes from observability data data so we spent a lot of time training models and working with the technology uh prior to founding flip I actually worked at AWS for about 10 years and uh about four and a half five of those years I was in the ainl uh Services organization um actually leading a team of Applied scientists that uh were building bespoke ml models for customers so over that four and a half and change years uh probably worked on several hundred poc's uh so I saw a lot of different AI use cases and uh a lot of um different Industries uh and uh part of what I'm going to talk about today are some of the challenges uh that that uh we saw with Enterprise adoption of AI and ml okay so generative AI is easily the most talked about Topic in technology today I I guess I can even start this off like who here hasn't heard anything about generative AI over the last six months right zero right everybody's talking about it uh clearly um there's something going on here uh really powerful technology but but we've been through these hype Cycles with AI before right and the reality is is that the majority of AI proof of Concepts fail to become production deployments and today we're going to talk a little bit about whether generative AI actually changes that right my experience at AWS between 2018 and 2022 uh we worked a lot with neural networks we worked a lot with sort of larger models um but it wasn't uh much generative AI it was still sort of the the previous generation of neural networks um and then even some classical machine learning so really the question is um why did these projects fail and uh is there uh anything we can learn to make sure that enterprises now can actually be more successful uh with the technology so couple lessons learned from my you know again fiveish years in the Ai and ml or get AWS so there's lots of enthusiasm to do poc's right we had absolutely no problem um maintaining a really full pipeline of of customers mostly Enterprise customers right in order to qualify for this particular program we offered at AWS customer had to be a certain size by Revenue um and there were tons of enthusiasm to actually work with us on building pcc's uh but most of these poc's failed to convert to production workloads so in other words we deliver the p customer would say hey that looks great but it just wouldn't make it into production and in fact more than 50% of these poc's failed to convert and I can't give you the exact numbers of the goals that we took uh in the years that I was leading this this part of the org but I'll just say that if we had hit 50% as the rate of conversion from POC to production we would have been over the moon right it would have been been just a statically uh you know good news so we have all these new technologies we have these managed llm Services is this going to be enough to sort of solve this problem to increase this success rate right we have open AI we have Amazon Bedrock we have cohere we have Google vertex AI um there's no doubt these are really really exciting new Services um and there's no doubt that these services do solve some of the challenges uh that uh we typically would see uh in in converting a a POC uh to production obviously the big one is the data labeling challenge right these generative AI uh models can actually learn from the text itself so you don't have that uh data uh labeling challenge which was one of the biggest obstacles for most customers to adopt machine learning if you don't have labeled data and you're doing supervised machine learning you're kind of out of luck but there were two other challenges as well uh there was uh the trust issue right so uh a lot of these previous generation machine learning models were very opaque even if we were were able to deliver a successful model to a customer the customer often times had trouble interpreting the results of the model or they they weren't sure why the model was returning one result or another uh and so as a result they didn't want to use it in production right it was still two black box still too opaque um we'll talk a little bit about the opacity problem and the trust problem in generative AI because it's hasn't gone away but it's a little different than what we saw before um and then finally there's the cost issue right so uh for many customers the cost of actually implementing a uh model in production was just too high right something like 90% of your costs when you're running uh an ml or an AI model are going to be an inference and uh in order to make your workload benefit from the AI if you have to do so many inferences that it's just going to balloon your your cost of running the model um that can be a deal breaker right um now the cost issue is interesting with these uh managed llm Services because they certainly do have have an economy of scale right uh and the reason they get that economy of scale is because the actual compute that's uh returning the responses is multi-tenant right so just want to be clear even in something like Amazon Bedrock where maybe you get a private virtual private Cloud uh endpoint into the service itself uh when you actually uh when the service goes to actually compute your response to generate the tokens for your response that's happening on shared infrastructure so uh you you kind of have solved the cost problem with these cloud services but there's a footnote to it right because for many Enterprises um that uh sort of Last Leg of isolation is is a deal breaker right they really don't want their data uh you know on a GPU that's next to like their competitors data and I know Amazon and and all the big cloud providers do a really good job at at segmenting and isolating data but again uh in the Enterprise space that can be a deal breaker so so for the rest of the the presentation we're really going to focus on those those those two obstacles the the cost obstacle and the trust obstacle and what we need to do to um overcome those in the age of llms just some numbers to kind of drive this point home right you'll see that 33% growth this is from a survey uh that Morgan Stanley did of CI in 2023 so so something like 33% of cios are saying you know I plan to have uh some kind of production generative AI llm workload in place by the second half of 2024 um and that's great you know that looks like a pretty significant acceleration you know that's one of the tallest uh bars on this this chart you know that that's wonderful news but remember the Po's that are going to determine which of those workloads Mak it into production are either going on right now or they're they're in planning right uh and so if you have a 50 plus% failure rate on conversion from POC to production workload load by the time we get to June of 2024 we all may be very disappointed uh if if if these are not converting at a higher rate right so the the estimates may actually be more um optimistic than than the actuals so the question is what can we do now what can we do if you're a a Founder an entrepreneur uh somebody who's building with large language models or even an Enterprise that's looking to adopt these models what are some of the things that we can do now uh to to hopefully improve the um rate of conversion from POC to production and what are some of the observations we can make um from the Enterprise customers that we're working with and that's what we'll explore in this session right we're going to explore um what we at flip AI are hearing uh from our Enterprise customers about llms we're going to talk about some of the challenges of using Enterprise data with these managed cloud-based llms um and then we're going to talk a little bit about how we're actually overcoming some of these challenges uh in the way that we're actually approaching building these llms uh at flip Okay so we've all seen these numbers um if you're in the same circles that I am on LinkedIn and uh Twitter or X or whatever it's called now um this is a slide that's been going around uh basically I've probably seen it you know randomly appear in my inbox or in my uh feed at least a dozen times in this last week this is actually taken from Sam Altman's keynote at the open AI Dev days last week uh and you know the numbers are nothing short of impressive right 2 million Act developers 92% of Fortune 500 companies 100 million weekly active users I spent a lot of time at AWS in meetings about um you know the the growth of certain AI Services there and I can tell you uh any product manager uh general manager would be salivating over these numbers right if you're a cloud services company this is really truly impressive but the numbers don't tell you everything right the um 92% of Fortune uh 500 for example it doesn't tell us whether or not those customers are really considering these Technologies for actual production work or if they're really still in exploratory mode right maybe maybe they're uh allowing their developers to open accounts that give you sort of an API key um and you can you know prototype stuff but you know it's still uh not likely that a lot of those uh types of efforts are going to actually wind up as production use cases so what are some of the observations that we can uh you know kind of call out uh to see if there's um ways to over some of the challenges uh in this in this adoption and and some of the things that we hear at flip uh when we talk to Enterprise customers and even though we're a seed stage startup it's kind of ironically most of our customers actually happen to be Enterprises um which is a little bit unusual especially compared to sort of the 2010s um kind of startup uh trajectory which is a lot of sort of product L growth um things have sort of change now uh and so we've talked to dozens and dozens of of uh Enterprises and we have a dozen or so that we're actually in active uh POC mode with and there's sort of three main uh requirements that that we've been able to distill from these you know dozens and dozens of conversations that we've had with with these Enterprise customers and the first is the models have to understand the customer's business data you can walk into a customer meeting with a really slick demo that uses some kind of public data set or maybe even some kind of sanitized or even synthetic dat data set that you've created to really show the power of these large language models and the first question you're going to get is wow that's really cool but I need this to work on my data and my data is really idiosyncratic because we've got you know Decades of business processes that we invented here uh and those business processes are what uh create this data or inform this data right so you have to be able to demonstrate that the llm solution is going to be able to derive legitimate insights from the customers uh business data second Point uh on this list should not be of any surprise to anybody that's working that's worked in machine learning over the last several years uh but it and it's really a binary requirement right there's only one right answer to the question when you get into a meeting with an Enterprise customer uh and they ask do you use our data to train models that you'll make available to other customers the answer is no you cannot do that um you have to meet uh this requirement um and In fairness all of the large um cloud-based llms as well as the hypers scalers that offer llm Solutions like adabs gcp and Azure they're going to have data processing agreements that that make these guarantees okay so that one you know that's kind of a a not so interesting point but um really you know you have to be clear that you're never going to use uh the customer's data to train a model that you make available to somebody else uh and then finally you know these Solutions need to be cost effective and they need to have measurable return on investment right um You you know if if if it doesn't in in the world that I'm in which is sort of Dev tools and observability it just becomes yet another um you know what 3040 $50 a month subscription that you're buying for your developers without any real uh understanding of how much leverage you're getting from that investment okay so let's go back to the trust situation right remember how I said in 2018 2019 2020 one of the biggest issues even if you were able to deliver a successful POC to a customer was how do I trust the results right um and back in those days a lot of times the output from the model was was actually something that was difficult for humans to understand it would be a you know a top 10 uh you know highest probability uh output and you sort of have to decide well what's the probability threshold that I'm going to basically assume anything above that threshold is is an accurate response you know human human brains are not wired to think probabilistically so that's just a hard one for people to get over um generative AI on the other hand gives you a response that's much more accessible to the human brain right it's text right you can read it hey this sounds pretty good this looks pretty good but you still need to know whether or not you know the the response is grounded in any kind of actual facts and if you're an Enterprise the facts are basically the data that you have that you believe to be true and to sort of be your ground truth right so if you have this model that's been trained on a huge Corpus of data and it's never seen any of your Enterprise data how do you make sure that the model generates a response that's factual um well you do it with something called retrieval augmented generation right so retrieval augmented generation is you sort of give the model uh a document or some facts uh when you ask the model to generate a response and that sort of nudges the model in the direction of kind of grounding its response in something that you know you know to be true um and this works really well and and in fact um most of the best practices and reference architectur you're going to see today around building an llm solution probably do something along these lines right uh and in fact some of the other things you might have noticed you know all this talk about Vector databases uh and these sort of new data uh bespoke data storage um uh systems that are that are basically you know being widely used uh When developing with llms that's directly related to this right the the vector databases give you a way of sort of retrieving those facts that are semantically similar to the um prompt that you're uh sending to the model so it works really well but you have a little bit of what I call uh the rag Paradox here okay so you need to First build this knowledge base that has what you consider to be those ground truth facts in it you may have all kinds of data in your corporate environment Google Drives with all sorts of documents CRM systems with case management uh you know whatever it is and you may largely trust that data you know you probably think yeah it's it's reasonable it's reasonably truthful um but using it in its sort of rawest form isn't going to be great right there's going to be some data curation and some data classification maybe some data enrichment you're going to need to do so building that knowledge base is is kind of a precursor you think okay great the llm should do this for me right uh I can I can feed it a whole bunch of documents I can ask it to summarize those documents for me I can ask it to classify maybe the topics that it sees in those documents I can capture all of that into some kind of um index like a say open search or elastic search index and I can build this you know really great uh tool for queering uh you know data or quering queering the documents that have the data that I that I need from them but you see the problem right if you use the llm to build that source of ground truth how do you know the llm is actually um correctly giving you responses that that you can trust right you don't have uh those facts that you can give it so that's a little bit of a chicken or an egg problem here um and so so you you quickly hit the wall um with uh retrieval augmented generation in in some cases not in all cases but certainly in some okay so how about training your own model right if it's going to be all this work to build this knowledge base just to do retrieval augmented generation why don't we just train our own models right hey you're a big Enterprise you've got big it budgets you may even have a data science team um you can partner with somebody like open AI this actually also was from Sam alman's um recent uh keynote uh and you'll notice though you know it's very expensive $2 to $3 million and several months of time just to train a custom model that is going to be out of reach for even the largest Enterprises um and keep in mind too that this is not a ontime cost this is actually a ongoing cost because models uh are trained on a certain Corpus of data certain knowledge and that knowledge gets stale over time and in fact Sam Alman even said when he was giving his ke ke note uh that one of the biggest complaints that uh they were getting from their 2 million developers using their apis was that turbo 3.5 which is the model that uh until recently was powering chat GPT uh had had a cut off date of 2021 that means it hadn't seen any data past 2021 um so that's a big problem right uh especially if you're querying for technical information um you know uh in fact chat gbt couldn't even really tell you things about chat gbt because it was trained on Data before chat gbt EX Ed right so you're going to have to keep training uh these these models over and over again and this is just a huge cost um and most customers aren't going to do this okay so where does that leave us right so um we had to solve for these problems right for our Enterprise customers and so we took a three pronged approach um the first thing that we did was we decided we were going to start with open source foundational models there's a lot of really good ones out there particularly some of the stuff coming out of Facebook we've had very good luck with llama and some of those um but we intentionally started with models that were not of the massive size right we we don't need hundreds of billions of parameters for what we're trying to do we only need say tens of billions um and what that means uh is uh you know the models are smaller right we don't have to show them nearly the same amount of data uh that we would to train one of these you know big multi-purpose uh models um we also spend a lot of time and probably the most significant investment we we make when we go to train a model is on that second uh item on the list which is um we pre-train instruction tune and fine-tune with highly curated data sets that we we produce that we acquire and produce um and these are drawn directly from the target domain that we're working in so this is largely observability data things like log files events metrics not so much because um llms aren't very good at math um but we'll we'll augment that with you know um issues like from from GitHub projects and Technical documentation API documentation right all of the things that developers are likely to look at when they're debugging a problem and we spend a lot of time on the data curation piece I've actually um recently had to to you know write some of the code uh that runs in our pipeline to to prep this data and and it's a it's it's an area where we actually um focus on getting everything right uh so that by the time you actually go to train the model which is going to cost you a lot of money and a lot of gpus um you know you can feel confident that that your model uh is going to work um generative AI uh conforms to the same Axiom as any other machine learning uh technology which is garbage in garbage out if if you train a model with really sort of ugly looking data it'll generate something for you but that something is not going to be very useful um and then the third one may be a little less sort of intuitive but what we do is we build the the inte Integrations uh that have the data that our model is going to actually interrogate um we we build direct Integrations into those systems right so in our case these are observability platforms we don't pull the data out we don't um index it we don't you know uh run embeddings on it and store it in Vector databases right we want to keep the The Source data as close as possible to the llm and and I'll explain a little uh bit more about that as we go through this okay so uh one of the things that we realized we could do uh to really come up with these really sort of good data sets for training the types of models uh that that we need to train is that um you know we realize that chaos engineering can actually work well for us so because we're training our model to identify um the root causes of various system and application faults we can actually run all of these different sample applications that use all kinds of different um you know uh architectures everything from sort of Monolithic kind of j2e type architectures that were really common back in the you know early 2000s to the sort of more modern microservices architectures that do everything asynchronously with Kafka uh message cues and and so on and so forth so we run all these apps that have a sort of lot of variation in the technological um or the Architectural Components um and then we inject fault scenarios into those right we have a pretty good understanding of what sort of the universe of failures are within these different Architectural Components and so we can create these chaos scenarios uh and inject those into these applications while capturing all the metrics events logs and Trace data uh and in in these very in these different platforms right so we get a really nice cross-section of of data uh and then we can use that to train our model so um obviously this won't work in every situation but this is an example I think of the kind of creative thinking that you need to be willing to to do uh to really get those data sets that are going to give you high accuracy models and in our case what we're really shooting for is that uh zero shot accuracy meaning can we bring a model into an Enterprise customer and show it some of their observability data that the model has never seen before and actually have it give meaningful um insights from that data okay so uh a couple of other things right um again going back to that you know 93% or 92% of Fortune 500 customers are using open AI I have no reason to doubt that Sam Altman isn't being completely truthful in his numbers but at the end of the day uh you know again what does that really mean right like if you build a compelling solution but the customer the Enterprise customer says you know what my security operations team uh really does not allow uh any multi-tenant type cloud service in our environment you know you sort of have to play by the Enterprise customers rules right if you're going to try to change uh the you know the the rules so that they they do something new it's going to take a what's already a long and and frankly rather painful procurement process and just make it any longer so what we've done is we've decided how can we remove friction from this process right how can we um you know basically get our technology into the Enterprise uh in a way that conforms to best practices and uh sort of security posturing that that they've already agreed to sort of uh work with and in our case the answer is you know single tenant deployments in the customer's own environment um and we can do that because we don't require massive arrays of gpus to run the models right I we can fit our models on two GPU instances uh which is well within the sort of reasonable amount of sort of compute uh cost that customers already used to um you know if you want to end a pre-sales meeting really quickly with a customer um about your llm solution um tell them that they've got to spend you know $80,000 a month on Nvidia gpus just to run your solution right that meeting will end right away um and and it's and by the way it's not likely that they have that capacity just sitting around somewhere in their Data Center and you know we've talked to some of the biggest most techsavvy uh sort of Fortune 500 companies out there and they don't have them right nobody can get gpus right now so smaller models work well for us um and then the other thing you'll notice in this is just those direct Integrations to the data um systems right so we don't if anything I would say draw your attention to what's not on this diagram right no sort of streaming data buses that uh require you know uh you know different Big Data technology um data engineering to sort of extract the data transform it um index it right we can we can basically query directly from the observability systems and it makes sense right these observability platforms are really at the end of the day Big Data platforms they have really nice query capabilities they index data already uh they provide metadata about the data we just use that right so we try to keep the the Integrations as simp simple as possible uh and in fact if you're using um some of the cloud-based services like elastic cloud or data dog you know we can integrate within minutes right all we really need is just an API key and an endpoint and readon access into those systems so this is this has been um I think a big uh sort of learning for us which is um you know you have to play by the Enterprises rules uh and no matter how cool a technology is uh that that's not enough on its own to get the Enterprise to to change their rules just for you okay so uh let's go back to the the truthfulness uh of the models right how do you how do you um convince your customers that your models are producing meaningful truthful outputs well in our case um you know whenever we generate what we call an RCA candidate and if you read that uh text down at the bottom um that's actually all generated by our language model but we always show corroborating metrics uh to to basically um as evidence that the hypothesis the model has produced is accurate right so uh you know on the right hand side we we we you'll see we call that supporting metrics and so if the model hypothesizes as it is in this case that the um issue is due to an unindexed database query what we will do is uh we will then ask the model OKAY given an unindexed query and given that we have all these database metrics over here which are the metrics that appear to be corroborating this uh hypothesis um and you can see we our our our models are even able to do some time series analysis and find Peaks and other things that that look like uh they corroborate the um the hypothesis of an unindexed query and then I know it's a little hard to see but under the flip AI logo there you'll actually see the AWS logo so we always surface the source of where the data comes from so in addition to the metrics themselves we always tell the customer oh and these are coming from your AWS cloudwatch metrics or whatever they happen to be could be data dog could be Splunk could be whatever and this is really really popular with our customers cuz if you think about what we're actually doing here it's really emulating the way humans think about a problem right if you're debugging an issue you're going to come up with a hypothesis hey I think this is due to an unindexed query and then you're going to look for other data that supports your hypothesis and that's really how we've we've um leverage these language models uh to really sort of mimic what the the human does uh but of course we can Surface you know large amounts of data uh without having a human to you know need to click through dashboards and you know run queries by hand right the models do all of that for us so let's return to the original question we posed at the beginning of this talk what problems should Enterprises actually trust an llm to solve right so first and foremost pick a problem that's rooted in specific domain data and and the reason you want to do that is is these are going to result in higher value Solutions right because they're going to involve specialized differentiated business processes right if you're picking a problem that's rooted in sort of very general data and there are business processes that that meet that criteria you can think of things like maybe um you know some standard sort of Human Resources processes um maybe some you know payroll type processes uh yeah sure um I mean a big you know trillion parameter model uh has probably seen enough examples of that data to be able to give you you know meaningful results especially if you use uh retrieval augmented generation but you're not going to really get a lot of competitive leverage out of that because all your competitors are going to very easily be able to automate those processes as well so you know focus on the things that really um allow you to get sort of high Roi of your your crown jewels which is your your business data um focus on problems where the domain data can be consumed directly by the llm um probably the second fastest way to um end a pre-sales meeting uh about your llm solution would be to um tell the customer they got on install a whole bunch of new Big Data platforms and bespoke um databases just to make use of your llm right they don't want to do that so uh if you can avoid having to do indexing and data enrichment and classification and vectorization and all that stuff um you know it it really will uh uh increase the probability of success um and then finally you know the third point this probably seems just glaringly obvious to everybody in the room but you know whenever you have a hyped up technology it's kind of easy sometimes to get lost in just the co coolness of the technology um it's very easy as a developer today to build a kind of cool gadgety uh app on top of an llm uh but if that's not solving real pain that's never going to really make it into the into the Enterprise um it's certainly fun to build those things on the weekends and on the evenings if you want to play around but uh just yeah avoid the temptation to think that just because you can sort of reimagine something that people can do today by using an M that that's that's going to be an actual product it's probably won't be right you really want to focus on things uh that are actually painful and that's all I have so thanks a lot uh appreciate the chance to speak and um visit us at our website and we'd love to hear from you