Devreal

SBTB 2023: Leah McGuire, Keynote: LLMs, what are they good for?

SBTB 2023: Leah McGuire, Keynote: LLMs, what are they good for?

Recording: SBTB 2023: Leah McGuire, Keynote: LLMs, what are they good for?

yes I've been uh doing sort of machine learning for a long time got out of out of science figuring that I uh actually didn't want to be an academic and wanted to make money and so went into the field of machine learning been doing machine learning at B2B company so really working on automl and delivering that for a long time uh and then uh basically went to pheros when I decided that maybe it is which is a uh platform for engineering operations in data platform for engineering operations information when I realized after boing around these various companies that maybe it is an interesting question whether uh why some engineering orgs are good at shipping things and some are less so and to do that I wanted to look at the data um so for this talk uh I should say that I have never been more concerned that I would show up to a conference and give the same talk as everyone else uh and I will say that there have been many great talks uh yesterday and probably today as well on llms and it has covered some of the same ground um but I will try and say it uh with some new insights and at least with my slide titles in the style of Terry pratchet as uh given to me by an llm and so um I will try to entertain you uh and the first thing that I want to talk about is really how we ended up here how it ended up being the case that I had to give a talk on llms because this is what we're doing if you talk to me in January this is not what I thought I would be working on at this point um and so what are good use cases for them why are we so obsessed with them and what are we doing with them uh and then I'll go through the actual use cases that we're doing at feros and how we evaluated them and decided that these were good use cases so the last couple of years as I'm sure you all aware have been just this explosion of foundational models which really changes the way people think about machine learning because you don't have to spend a year developing a model to do a particular task you can just sort of grab this off the shelf thing and it will give you an answer so you know we're either in the future or in your favorite uh fantasy where you can just walk up to hex the supercomputer and it will give you a written language answer to any question you can ask or an out of cheese a whichever whichever is the right thing to do and so we've gotten to this point where just everyone is working to integrate llms into their software and I was talking to companies this summer and it's just across the board everyone wants llms and everything and so uh the first iteration of this was just with text to image right and so this is fun you can generate fun images but it was really more of a party trick right like go to a party and you're like can we make this thing uh produce an image and you know you laugh about it uh it's not this transformational thing for most tech companies um it certainly had a big impact on Creative communities and there was a lot of debate of like what's the role of humans um does this does this sort of make them obsolete for a lot of these jobs we got great memes out of this which is fun um but I think the the conclusion was more along the lines of no you still need humans to get good creative workout to get the right thing you want humans to do it so it's really a tool um and it didn't affect me nearly as much uh but we get to llms and you know everyone's just crazy about them like why is it that if you can put in natural language everyone wants to do this and I think it's really just because language is so important to humans right it's one of the things things that makes humans uh what they are and it's been defined um as sort of this this stepping stone of like other animals don't have this so there's a lot of work in neurobiology showing that other animals do have very similar things um they can vocally learn they can do these things but uh linguists sort of go back and then redefine what language means to say no no it's this complex ability to Define any abstract comp uh concept with words and that's what makes it human I think this is a little bit the case of defining a metric to get the answer you want uh which is also something that seems to happen with emergent properties in llms but nevertheless it's it's hard to overstate how important it is we wouldn't be here at this conference talking about all of this fantastic technology if we hadn't figured out how to write down information about that and share it with lots of different people and learn on top of it and so this really does seem like a building block um and if you haven't read it uh I I highly recommend the blog from Step wfam what is chat GPT doing and why does it work uh it gives a great explanation of sort of how these models are functioning uh and sort of highlights the difference between uh being able to produce language and being truly intelligent right and and I'm going to spend several slides on this uh just because I've seen it so much in the news and it kind of IRS me so uh there's this idea that these language models are just predicting the next token and it's it's predicting the next token based on learning from a vast amount of data right more written language than humans could really take in and learn from um but it's still predicting the next token and I just want to say that you know the word Apple to me is a lot more than the next token it's a lot more than the word before tree or sauce it's thought and feeling and a taste and a memory memory of apple orchards and I think that's really important when we worry about sanant and intelligence and what these machines will do uh and just to go a little bit further than that I'm going to pull on Neuroscience just a little bit uh so this is a figure from a paper called language and thought are not the same thing uh and what this is showing is actually uh two areas of the brain that are involved in language and this is language comprehension and language production at this higher level uh of saying like you can convey information you can understand it um and the uh the areas not highlighted in red are actually other language areas that are not this higher lever so they're motor production of language seeing and understanding language so perception uh these are excluded from the study and what this shows in both functional Imaging and studies of people who have had Strokes that damaged these areas is that you can do other important thoughtful things without these areas without language uh and so this this chart uh next to it basically shows that uh during language uh here um you can see that these areas are highly active for sentence production and things that sound like words but for doing things like math which requires a great deal of thought or doing things like working memory memory so remembering one thing after another like Donald Trump can do or uh basically cognitive control so navigating through space using definition of things rather than geometry or music these areas are not involved in that and in fact you can still do all of these things without those areas so there really is more going on in human thought than just predicting the next token which isn't immediately obvious like it's hard for me to say that what I'm doing right now is not just putting the next token saying the next word after uh what I came up with before right but it really is different um and if you look at the complexity these models are you know very very big uh I'm going with llama to here I refuse to speculate on the size of architectures that are not open- sourced but uh this is a lot of parameters it's a lot of weights to fit and it's trained both on just the vastness of everything that we've written and reinforc to produce answers that are good for humans that we' like to see so it sounds right um but again nature is more complicated so if we just take a little stroll through Wikipedia and look at the brains of various animals uh llama smarter than a sponge right there's no neurons in a sponge it just sort of exists um but you don't have to go very far to get much more complex so you can see honey bees actually are quite complex animals they have about a million neurons they have about a billion synapses which is this equivalent of uh weights in these so it's the connection between neurons uh I'm going to leave aside the fact that animal neurons are orders of magnitude more complex than the neurons in these models um but you can go just a little bit further and see that a mouse sort of blows these models out of the water and so my point here is that maybe we should just spend the same amount of time worrying about these models being scient as we worry about mice being scient which is not to say that mice are not cool and we shouldn't spend some time worrying about it they can vocalize they might be doing a giant experiment expent on us to see whether or not we can they can figure out the meaning to life the universe and everything but but it's still not something most of us spend a great deal of time on it's not the same level of news Cycles so so we're back to this idea of talk uh because these models can talk to us we feel like we can use them for anything and I have seen many proposals of people trying to use them for everything uh and they're a fantastic multi-tool in many ways um but a multi-tool isn't always what you want right like it's a it's cool it's fun it's an interesting place to start but if you're going to do a serious woodworking project you probably don't want this AI generated Contraption for it you probably want some real tools uh and I think that's maybe the lesson is like you can get started with this but it's maybe not the end goal um and so one of the use cases that I've seen uh permeating this is just this idea that you use these models to help you write uh articles and information documents I was super excited when they came out I'm like I'll never have to write another tech technical document I can just write an outline and then the llm will fill it in for me uh but if you've done this you'll realize that it tends to fill it in with very generic information and not necessarily make your points uh and so I did just a little experiment to do this um saying we if we get to the point where we're we're having these LMS write our information which many people are and no one has time to read all of that generated information so of course we're using llms to summarize this as well and then you use those summaries to generate new uh documents sort of what happens to the information right and so I did this I took the uh outline for this talk that I showed you at the beginning uh and I fed it into an LM I said write me an article and then another LM summarize it another llm write me an article and then I told the final llm to summarize it for me and what I got back was still on llms which was really impressive I wasn't sure that was going to happen but it has no resemblance to what I fed in it seems like something that was probably uh specifically trained into these models to return as the way the companies that Supply them want llms to be talked about um it's it's just sort of a very different thing and I think this is an important piece to keep in mind as this is sort of permeating the information that we see in the world um and again like the the use cases that we have the use cases that people want to do there's uh a lot of different dangers and different uh harms that could come from them and so there was a great uh image chart from Alexander uh toov early on when these came out and it's just this question of like when is it appropriate to use an llm uh and the first question very importantly is do you care if the answer is right if you don't go for it but like if you maybe care then you should consider whether or not you can tell if the answer is right and this is really tricky because it is hard to tell sometimes the answer is right the LMS produce very convincing uh hallucinations or just like mistakes things that look like they would be produced um because we've trained them on all of the things we've produced all of the garbage and all of the actual information uh and if you're not willing to if you can't check it and you can't uh assume responsibility for it being wrong then don't use it um and I think that this is uh increasingly important and even if you can take responsibility for this even if you don't care if the answer is right so this is an example of someone just goofing this is a real screenshot from my phone someone just goofing around with an llm trying to get it to say something funny uh and it became a sensation on the internet and then it became the informational blob on Google for uh the search of what do you uh countries in Africa that start with k um and then you get this like seemingly informational piece of thing that says no no there's no countries in Africa that start with k the closest one is Kenya which starts with a k sound but is actually spelled with a k sound like this is obviously wrong but there are there are many examples of things that are less obviously wrong uh and so you need to be aware that this is already happening this is already sort of polluting our information space um and then of course the the question of uh these are really really big models probably unnecessarily big models for many of the tasks that they're being used for and people are just sort of sitting around talking to them playing with them and filling the atmosphere with with carbon um so there's there's many harms here there's a great paper out of Deep Mind that lists these as well as all of the sort of maliciously harmful things that people might do uh I'm not going to focus on those so much as the things that you know we could accidentally do because I think we're we're not trying to to do harm with these models here uh but it's still possible to carelessly decide to use them for the wrong thing so what are good things to do with these models what are good jobs uh for llms and again language is a really powerful and natural interface for users it's much more natural than than the whole like web computer interface that you sort of have learned as a child so you're very familiar with it but it's not the thing that humans have evolved for millions of years to do and so being able to interact on that level and put this information into language uh becomes really powerful and so the things that we decided to try first at feros uh were two things uh basically to make data access easier so the first is just putting captions on charts so a picture may be worth a thousand words uh but it really helps to have some words to explain what the picture is about and so we wanted to automatically capture that and give people an aid in viewing sort of a complex dashboard that they may not have experience with and and really understand it and then of course there's um this thing of it's actually really hard to formulate a good question given a brand new database and schema and make sure that you have the right answer and so we wanted to provide an aid for that and the way that we did this um I'm going to talk about query Helper and not so much about chart explainer for this um but we we focused mostly on using llms as an aid and I should say that I joined phos fairly recently and so much of the work that I'm going to talk about was not from me uh I was it was like my dream job of showing up and being an intern other people had already thought of the project I just had to implement it was great how I want to retire basically just do it don't have to think about about what you're doing um but much of the work was done by my uh three times colleague uh Ken C shown here uh as well as a lot of work uh from Sarah Asher uh Sarah Asher Natalie Casey uh Shuba Nar and Chris rupley uh in terms of just defining what what should be done so I want to give them the credit for actually thinking about these projects but you know data is hard this is the real screenshot of the schema at feros uh and you know it looks terrible uh but it's not not it's not actually a bad schema it's only 58 tables uh and it's well laid out it's got like good connections uh it's it's standardized across all the customers um but if you've ever started at a new company and had to figure out where the data lives and uh how to connect it and what to do with it it's hard right it takes a lot of time uh to learn and it is easy to make mistakes and make mistakes that really give you the wrong answer in a critical way and so we wanted to make it easy to approach this and have our customers not have to spend a ton of time figuring out how to use the product uh and so in order to do that we wanted to make a helper so this is what the UI looks like this is the final product that we made uh and what it is is the ability to type in your question about what you're doing and then it will we will go and return related uh charts related visualizations and summaries that have already been made within your piece uh and we improved this using llms as well so we improved upon the basic search uh and then we also create a textual how-to guide so this is how you're going to get your data this is what you're going to do with it and this will answer your question uh and then importantly we parsed out the table information from this llm generated how-to guide and enriched it with handwritten instructions and information about each table to say this is how this table should be used this is the join keys so that the user can actually look at the instructions verify with the information about the tables and just sort of quickly understand what they're trying to do so putting some of the ability back on humans because we're really trying to build a tool to make it easier for humans not have this replace the humans and I'll talk a little bit about whether or not that should be the end goal um but of course in building this the first thing you have to ask is are we doing it well enough to ship it in front of people so uh please don't try and read this but this is just a picture from a review article of all the different ways that people try to evaluate llms uh and they do this basically uh because they're trying to evaluate them for everything right so you're trying to evaluate whether or not this llm can pass the lsats and write SQL and write code and do all of these different things and I don't actually need the LM to do everything it's I it's cool that it can but that's not what I'm trying to do right I'm trying to do a specific task uh and so there's other ways of evaluating such as the hugging face open llm leaderboard which allows you to look at various models performance on particular benchmarks for particular tasks and you can use this to pick your model which is great it's a good start it's still not your task it's still not evaluating that uh which is tricky and it's become sort of more and more common for people to just say well I can't figure out how to evaluate this I'm going to feed it back into the llm and ask the llm how it did uh I think we can do slight better than this and there are a lot of problems with this not the least of which has is that these models have been shown to prefer their own outputs uh and so so they're more likely to judge themselves as good uh above humans above other models and so I think it's it's a little bit too much feeding pigs bacon to really uh provide a good uh answer here so how are we going to do this we're going to focus on what matters does the spelling in your AI generated picture matter or just the fact that it looks like what you're trying to do uh we defined particular metrics that we wanted to do and kept in mind uh good heart's law which is that the metrics are not the actual goals so we refer back to the metrics look at the actual data and evaluate based on that to make sure that they're really capturing what we're trying to do and not just something that we're optimizing for uh and so the metrics that we decided to use were Rouge one uh and jakar similari and the Rouge is a classic summarization metric it's basically looking at engrams in our case we just looked at at single words uh single engrams um for a uh for the summary and a reference um and the overlap between those uh you can do this as the classic uh definition is of recall uh we used a library that also calculated Precision in F1 which is the harmonic mean of precision and recall uh and we chose to focus on F1 so basically the similarity between your reference uh explanation and a generated explanation uh and then it was really important to us the tables that were recommended be correct this was sort of the most important part can it find the right tables to do and so uh we used to card similarity so the intersection of the tables in the reference and the generated response over the Union in order to do that and we did this for both the individual table names as well as all the the tables plus fields that were done uh and so the first the first question is where do you get these references to compare your llm results to and what we did is look at our uh dashboards that are written by uh Us by the engineers and data scientists at feros and use those as examples of correct usage of the schema and so uh each of these these uh dashboards contains many charts if you take take the title of the chart as the question and then look at the underlying information used to generate the chart you can provide an explanation for how you got that chart and so we did that and we did a number of iterations of how this was done both by hand and using llms which I'll go into uh in later slides and found that it's really important to have a lot of explanations and to have really good explanations and the explanations look something like this so basically for for a question question time to resolve incidents uh we have a primary data source the join information uh the filter information uh the breakout information and finally uh any aggregation and you'll notice that this is very similar to SQL but we do not directly expose SQL to our customers and so this is just an an explanation of how to use a UI in order to get the information um so the first thing is like how do how did this work how do we uh get the right prompt can we get an llm to perform well um and so this is the information based on tables so did it produce the right tables uh and we did this in a number of different considerations so uh with no examples so zero shot we just gave it some tables and asked can you tell me which tables I should use to answer this question uh and it did okay uh if you gave it examples just particular examples it actually did much worse uh and so if you give it static examples that don't have anything to do with the question it will sort of over index on those examples and give you answers related to the examples and not the question but if you give it relevant examples it does much much better uh and more examples or more tables don't actually help this in help this um so looking at uh in addition to the tables you can look at the SK schemas so the field information that it does uh what we found is that without examples it actually failed to put things in the correct format and so we couldn't parse out this the field information uh you needed some examples to get the right formatting for this um but again like more examples relevant examples was really the way to get better performance and the same with the format so the overall format of the explanation was better as you went to more examples uh better examples and so we spent some time really thinking about like ah so in rag it's not about the size of your prompt it's not about stuffing all of the information in it's about getting the right information just the right information into your prompt so that you can answer the question uh and this this is really critical and so in order to get the right information we actually looked at how to generate these examples uh we used llms to do to do the generation uh in a number of different ways uh we used manual explanations uh we generated it from the underlying SQL that accesses the data uh from the from the Json uh that generates the data and finally um we went in and manually fixed the Json information in order to get better results so it was this this combination of you can use llms to bootstrap your data but you're still going to have to spend time making sure that that's quality data because that is the most important thing is really getting good data um and so all right I'm gonna I'm gonna give this last piece of new information and then I'm going to hurry on to the end because I've way over spent my time uh so uh the last piece was uh picking which model to use and so we started out with open Ai and had a lot of issues with timeouts and um basically token limits and outages and it was just really really slow and so we moved to trying AWS Bedrock uh with Claude V1 and V2 and found that it was much more reliable it turns out AWS is good at being a platform uh and so so we did this and then we looked at the overall quality this is one example but there there were many others uh and found that there's not really a significant difference in the quality across models if you squint maybe you could decide that you want a GPT 4 to be better here but it was not significant it was not visual you need to do like many thousand more uh trials in order to get a real difference and so it was certainly not enough to make up for this nonsense and so we went with uh anthropic through bedrock so in order to build this we did a number of things uh so we used llms to basically introduce semantic search into the related chart and then rank the information uh we used prompt engineering to get as much relevant information as possible into our llm query and and really make sure that that query was performing well and we then parsed that in order to give additional details about the data so that the customer can make the right decision in how to use it and so the end point of this was that we got something very very quickly that we could ship in front of customers and see if they were interested in it see if they cared and what this means is that we didn't have to spend a year building a model to see whether or not it was worth doing we can put this out we can monitor whether people respond to it and if they care then we can iterate then we can make it better we can see uh whether or not we need to f- tune we can see whether or not it's worth the investment of doing models whether or not we can ever get to the point where we can automatically generate the charts and you know I think the answer to that is not with the current iterations of these mod models um but maybe someday maybe with something more specialized that's not sort of this this multi-tool that can also do the lsats but something that's really specific to SQL like we saw in in some of the talks yesterday uh we'll be able to get to the point where we can trust it to do this well um so we're in the future you can just ask a model to give you a holck and it will give you something that looks cool and is vaguely related to the thing you were thinking about uh and so it really is powerful but it's not quite to the point where we need to lose our minds and I think that's sort of the the lesson that we took from here is that if you take this multitool and think of it as just like a new way to quickly iterate on shipping an AI product and make sure that you evaluate it make sure that you have a good business goal uh it's really quite fun and Powerful but not necessarily something we need to lose our minds over and you should still make sure that you have a good business goal in in using llm uh so thank you very much