Devreal

SBTB 2023: Mary Grygleski, Boost LLMs with Event Streaming and Retrieval Augmented Generation

SBTB 2023: Mary Grygleski, Boost LLMs with Event Streaming and Retrieval Augmented Generation

Recording: SBTB 2023: Mary Grygleski, Boost LLMs with Event Streaming and Retrieval Augmented Generation

okay welcome everybody who's here uh for my talk uh you have two other better maybe Alternatives but you chose to come here so thank you very much okay so I'm going I'm here to talk about how you can boost your llms uh with event streaming and retrieval augmented generation but first of all may I ask to how many of you you are already working on gen types of uh apps maybe a few okay yeah quite a few so um and actually also let me ask to how many of you are working with let's say Vector database too like maybe yeah just a few okay that's that's cool I just am curious because I've been also like traveling as a developer Advocate myself I just travel to different places I just want to get a feel like of each region of the world um there are different uh um like you know kind of adoption of geni and it's very interesting because as we know gen chat GPT is now taking the World by storm and uh there just a lot of talks a lot of things there are many things you know around this so will be interesting too okay so I'm going to start so let me this is the agenda so I'll I'll just first um talk about who I am first very briefly and then also because it's a 30 minutes thing so I have to be quite uh f fast and brief too so pardon me if I may be skipping over something or not explaining something but um I will give you my contact information and also company um my company information all the products that you can also you know always come to us too so this will be sort of an introduction and hopefully we'll um you know pique your interest if you are curious about it um and then if you are here already heard about it I pardon me if it's kind of a little too slow for you but I have to assume to kind of lowest common denominator maybe not everybody is familiar with it so okay so just a brief introduction to that I'll be bringing you know bringing forth and then also discuss then why we want to get into you know uh boosting llms what are some of the drawbacks and also the options for improvement and then also with rack how rack can help um and also wanting to introduce to you an open source library that my company is working on call lstream so and also a quick demo I hope I have enough time to also show you something fun so that's a the plan who who am I first of all I'm Mary gesi and I'm a senior developer Advocate at data Stacks data stack is a company here based here in the Bay Area in Santa Clara how many of you actually have heard of data Stacks are familiar okay so quite a few how many of you are using Cassandra by any chance okay yeah so great and also let me do a bit of promo here Cassandra Summit in December so I also have promo code so I do want to um you know encourage all of you to go not just about Cassandra but there's also AI defa uh collocated too so um so that's kind of exciting we're kind of getting up you know ready for that too so okay so data Stacks so I actually joined data Stacks real quickly last year as a event streaming Advocate because we um have adopted uh Apache PSAR so we are very much an open source company Cassandra and then also Pulsar um and we also have a managed Cloud platform called estra so that's something I want to share with you as well uh later and then okay so that's about my company that I'm a developer Advocate there as you know as developer Advocate we travel everywhere and then bring goodness to to all developers we want you to succeed we want to share with you all the knowledge as well anyway so here's my background I'm a Java Champion I've been actually working with Java the most you know among other languages I'm also the president of the Chicago Java users group and these are just some of the things I been working with the poar uh also I was at IBM previously as an advocate uh for three and a half years so some of the open Liberty uh open source stuff so I've been like with open source for quite some time and then about maybe four months ago my company decided everybody's going to work on gen and we have our Vector uh database uh actually it's Cassandra we added in Vector data type and also some other goodies to which I'll share with you in a little bit as well so um we' like to encourage you too if you're looking into gen Solutions looking to our company what we can offer for you to and these are just my uh contact information again I'll share with you towards the end of this presentation okay so I'm going to just start first brief introduction to geni and Chad GPT as we know C CAD GPT came out right the version 3.5 that shocked the world was like in a year ago November of 20122 so since then too everybody's been like excited and a lot of confusion as well so here too um let's kind of take a look at what it is in my case too because my company said let's work on it so I started looking into it and as such too when I started working with AI prior to that too I didn't have as much of an AI even though I did some work with it actually with streaming with reactive streams uh with other speakers too on uh machine learning reinforcement so I just thought this is a great way for me to get back to doing some really formally into AI so and I found out gen AI is is C different and it's bringing a lot of excitement but first of all too let's take a look brief introduction to gen AI essentially too if you kind of think deep it's really about automation we don't want to do the work human beings and we want actually machines to do it for us and that's the ultimate goal and this actually also again for those of you maybe if you're are somewhat new to this field and it involves a lot of data because only with data that we can actually you know kind of bring out what we're trying to look for right data is of utmost important in the AI field and as such too this is like the famous diagram like describing like an onion describing ai ai essentially a mimic intelligence or behavior pattern of human beings and you know or or any other living entity right and then machine learnings kind of if you kind of get deeper into it it's basically deals with techniques in which the computer can learn by you know from data and not so much about us kind of writing code so to speak but it's really you want to train the the computer to think in that case and then actually further deeper down is basically you want really to inject intelligence into it and that's when deep learning is you know kind of in the core of it the neural networking that really mimics how our brains the neuron neurons kind of work too so that's a real quick intro so fascinating look you know into the Gen AI era so what is Gen AI right so it's a disruptive field in AI it has the potential to change the way we create content and consume content and basically it generates contents and it's very important is will be based on prompts right prompts is what you actually guide what you want to try to search for and then you know have the bot have this um AI the the artificial intelligence bot to kind of search for answers for you and it generates contents so it uses compination of machine learning and deep learning to produce the contents um the thing is though with geni and before geni was actually more predictive AI so it's not really that new in some sense however predictive AI is really dealing more with analyzing data um and also the purpose of such is perhaps making more business forecast weather forecast so as you can see it's a bit more kind of businesslike so to speak but it also has a time kind of like valid or expiration to it because if you're making business forecast you usually say a five years forecast and you can of based on some data you generate some kind of you know um predictions on it um but the thing is though five years after or even just one year in you know today's world it will be expiring you know it doesn't become valid anymore um however or even like another example would be weather forecast to you can use predictive AI um but the thing is the Gen AI is kind of disruptive is because now as we all know you basically are working with kind of things on more of the creative side you can have it right an essay um you know generate images um and actually the more um kind of uh you know kind of more desirable thing would be like the multimodality in which you can actually have text and telling it some stuff and have it find the images that you need or generate whatever it is so and then as such too it can also be dangerous too as we know because you can actually kind of tell it to do something and it goes and does maybe put your image image on and kind of pair it with some big you know you know kind of criminal kind of person and then you could be incriminated because of that so stuff like that so well I don't need to describe too much now so let's take a look then into GPT right uh generative pre-trained Transformer so essentially that's what it is kind of mechanism to kind of make it happen so essentially GPT will take simple prompts right in human language and that's what also the difference is that now we can actually give inputs like we talk to one another no longer is it restricted to how we used to be you know are still we are programming there are certain rules you need to follow if you don't follow that it won't produce give you the results right so that's the thing too GPT is that you can take human conversational languages as input and then have it do pattern matching or similarity searches and then return you the results the responses too so that's what is fascinating is that now it can produce contents like writing an essays uh writing a piece of music mus and generating code for us and most of us here I believe you're a developer so that's pretty fascinating so to me I find that geni definitely has its value it's not going to take away our jobs but it's probably going to take away a lot of the mundane things if we can kind of train it to do all these mundane things that you know that then then we can free up to be doing things that are more creative is what I'm looking at okay so this is just a really quick kind of History too so it only started up you know GPT it you know in like 2018 too that that's when Alec Redford's paper on GPT came out of a language model is published on open AI website so the the reason I'm bringing this up is that is this really is come came on fast it's like you go you know going in the ocean and the waves coming and it's basically oh you see it coming and it's here you know within five years and the the waves keep keep coming too so as you can see then GPT came out 2018 then 2019 was a gpt2 and then 200 you know fast forward too a bit 2022 when stability AI developed stable diffusion so that also kind of immediately change the whole kind of uh trajectory too because it now allows you to do text to image type of model from there and from there that's when Dolly came out and mid Journey For example and then very soon too everything kind of happened like really like the waves come you know stronger and stronger and in 2022 just a year ago November 30th that's when GPT 35 got released and within 5 days it reached 1 million users so it's very disruptive too so okay so without going all the details we know that this whole waves is contined to kind of come and also too there are things that we have yet to discover I think that's what the exciting thing is and that's what we are all on the journey you know to kind of work work towards it and really like we're discovering new things as we go okay so now a bit of a kind of more the kind of basic terminology NLP so NLP is a interdis disciplinary subfield of linguistics of computer science as such it relies a lot you know on the input is dealing with languages so we need some kind of you know kind of way kind of studies or kind of a model kind of discipline to kind of guide us so that's when L NLP came out and then it also is basically helps to process n natural language data sets right those huge data sets all of these things and it uses like rule-based probabilistic machine learning type of approaches too so it enables computer to learn from contents to including the contextual nuances of the language itself and so so the the thing is that with you know generative AI that you're not just like asking it to do something but that we're expecting it to be more intelligent and draw insights from all of the documents that it can find to so now then it leads to like llms large language models and that's what I think a lot of developers we're working on these days are the llms that's what we interact with it is a type of machine learning model the foundational model um but the thing is that is you know in general too we don't kind of do kind of pre-training on it because it's just very expensive and it's huge amounts of data and it takes days even you know kind of like taking many gpus it's going to take days or weeks you know even months maybe sometimes to train it um so that's why there are like large corporations that are coming forming consortiums open AI for example that will generate this or pre-train all these large language models so the idea is that LMS is what we kind of interact with there are a apis that kind of allows you to interact with it as well it answers questions like from the promps you know like a human being it analyz the sentiments and kind of really good for chatbot type of conversations that would be one example um and then some examples of the API and Frameworks that we use today for this uh will be L chain uh llama to um and also uh Palm which is from Google and then hugging face as well among many other libraries too but these are kind of like the major ones and then me being a more of a Java person if you're a Java person then there's also for example Jay Lama is actually one of my our chief Architects at data stxs he developed J Lama is a Java Port of the Llama library and then also there's a j Vector too it's our founder of data data sex he designed that and is works with project panel if you're familiar again with Java and essentially it supports like simd type of um operations too so it enables like vector search to be super fast and also deals with the storage too so I won't get into all the details but these are the links and then there's also lanching also has a port to uh basically it's in Java it's called Lang chain 4J and also llama 2 which is a port of llama C so these are just examples if you want to start experimenting with it okay so now let me then also uh before getting into rack to Let's also understand what Vector database is and that's also one of the you know kind of primary things that my company is working on and basically uh Vector database is a purpose built database right that serves up Vector data type you know so the purpose of it is to handle complex machine learning uh type of searches and all of these things and you may also have heard of too there it works with Vector embedding so essentially if I kind of describe to you vector embeddings is essentially it takes maybe like a prom of strings that comes in and it tokenizes it breaks it all up and basically translate this you know your tokens each string it translates it to a numerical representation of it so in other words if you kind of look into a vector database a table if you do a select on a vector data type it will come back with an array of floating Point numbers for you so they kind of represent um all these strings that that you want to kind of insert into your database so the purpose of such is that it's being stored very efficiently in that uh format as well as for searches and queries too and it's basically an automatic feature engineering um type of uh kind of database and also if you kind of ask me what type of mechanism it use right the algorithm it uses for example in Cassandra the vector database or we call it estra DB uh Vector is basically uses this approximate nearest neighbor like Ann type of mechanism to do the proxim searches too so okay so now I brought up the again going back to Vector embeddings um what is it being used for right so essentially is used for searches right so the results to are Rank by relevant to some querry string too that's kind of that's what it's for primarily then you can also use it for clustering as well as like recommendations um I'm sure if you've done shopping and you do all of the recommendations you can use that uh Vector embeddings uh type of data to do recommendations too and also anomaly detection diversity measurement and classification as well so these are just some of the examples of how you can make use of vector embeddings okay so then let's kind of go back a little bit there's Vector database and I got asked the questions why can't we use traditional database so database such as postgress a lot of folks are using and I understand postgress also has Vector database uh capability I think it's called PG Vector um all of these things um but the thing is so let's take a look at traditional database the thing is you can still use traditional relational database however we're talking about lots of data and how do you deal with it right so it would take a much longer time the performance will suffer too if you kind of use traditional uh relational database the reason is that in machine learning gen AI type of scenario the data will be best kind of being you know stored and also being searched upon and and everything um you know be because the thing is that we're dealing with data that we want to identify patterns and the dimensions and the relationships so they basically too using Vector data type is the the fast way like if you are kind of math major you understand it is linear algebra behind the scenes using idian searches um and also cosign similarity kind of searches behind the scenes so actually I have to kind of say it was funny I I was actually a math and computer science major that many years ago and I think is I remember back then in college I was like oh what do we use it for I never could understand Vector math right and Matrix and all of these things but now I do I finally have a job that actually is making use of my linear algebra kind of skills in it even though not every day but it helps actually helps me a lot to immediately understand how linear algebra is being applied so this is an area so okay so so that's that um and then now um let's kind of real quickly then talk about drawbacks of lrms and this is the title of the talk is boost your lrms right with r so let's take a look too llms you know what are the uh kind of you know kind of some drawbacks of It kind of if you look into you know the the barrier of AI as we know um in traditional machine learning if you look at it well Machine learning is always you get the training largely broadly kind of divided up you know the steps are training phase and also inference phase so as you can see there's training data coming in there's machine learning engineering It produced a model the model is then used you know to kind of uh use for inference is basically that's when you you kind of use it and apply you know the the you know apply techniques to kind of search based on the model for what you're looking for like that so that's kind of traditional ml the thing is though if you kind of look into this too the training phase will cost a lot of money too as we kind of already understand right usually we're dealing with data it's come in large corpor of data comes and you need to gather the data clean it store it process all of these things and and it's also hard to find people enough people data scientists to kind of understand and work on it too the process itself is also very long and uh and basically you know not reusable it's not just very efficient you know that's the thing so now comes gen then that's when we talk about these large llms uh for for GPT type of usages same thing you get you know on the training phase you're still training all of this and then now with geni is that then we have the input of the prompts that coming in and and then inference space we also Alo involve the chat completion uh kind of phase and then generate you the next tokens and words all these things and then if you kind of uh to take a look into it too um basically too the llms itself also has a lot of um shortcomings too so let's take a real quick look into it so llms are basically parametric it can't memorize all of the knowledge right in their parameters too it's kind of has limitations and they're also probabilistic too because for example in example like here it's like capital of Texas is it can return you a whole bunch of things based on all of the measurements um of the results too and then also too um the another shortcomings of llms are that it is opaque too meaning that it's not really possible to interpret explain or add some annotations uh to its answers too it's pretty much kind of static and also training lmms is expensive as we all know right it's it's also not possible to add or remove knowledge from it it's kind of really a chunk of things is there that's all you have but what about right if it's uh think of some ideas to improve you know um you can always like create your own Foundation model but the thing is in reality it's not possible because it's so expensive and uh that's just not possible for anyone to do it you know unless you're really a rich richest guy maybe even he can't do it and you can also think of like fine-tuning uh and retaining an A llm on new data however with fine tuning tooo it also has its um shortcomings too it's not really recommended um in fact not the you know kind of good way and then of course there's also prom engineering too is basically you want to tweak you know your input to Cox the output um so that's that's what it it is but let's take a look then into uh prompting too so using gen is also about like building the prompts too um so the problem to solve you know is no longer in the model but it's in the prompt to however if you kind of take a look into prompting too it's basically also uh the size of the prompt is limited as well as you know right it it it you are still working with against the same corpor of llm data in there so it may not be you know ideal too okay so and then and then actually let me kind of skip that part because I want to still show a quick example too but that one is kind of repeating an anatomy of a prompt so to kind of summarize is that the problem is that the domain specific data is really huge too so you can be training an llm or let's say the llm is taken from 20122 then what happened between the time of 2023 and now you get that whole chunk of data not in the llm then what you what how do you kind of handle it right so this is when like our the the rack pattern the retrieval augmentation augmented actually generation pattern can become useful in here let's take a look then in what to see what is rack right so rack is basically a hybrid framework that integrates two components of these rack models one is the retrieval model and the other is the generative model so the the purpose is that you want to be able to produce text that's not only contextually accurate but also information Rich you know that's the idea so let's take a quick look at the retrieval model in a nutshell right the retrieval model acts like a specialized kind of librarian so it pulls in all the relevant information from a database or purpose of documents too and this information is that fed into the generative model so the generative model think of it is like a writer it takes all of these you know data that's been retrieved on the retrieval side and it crafts it into some coherent and informative text you know based on their retriev data too so basically the rag model involves these two kind of work they work in tendem to kind of provide answers to that we are hoping to be like you know more accurate and also contextually Rich too using this model so now and I just want to kind of quickly kind of squeeze in this is that the mechanism that's basically use you know in um doing like vector data in a rank model you could be also be kind of having data that's kind of you know you you're taking it and then also inserting into the database Vector database as well so this kind of quickly explained but I realized I don't have too much time so I won't explain too much but I'll share with you if you want to take a look at these um you know kind of slides cuz normally I do this talk like 45 minutes so now I kind of have to kind of skim over but I want to leave you still with these slides if you want to look at okay so these are dealing with Vector search and Vector store to in that but let's kind of take then quickly into look into the generative model generative model is basically the role is again like as creative writer you synthesize the retrieve information into coherent and contextually relevant context uh into the text right and it's built upon on the llms too and it has the capability to create text that is grammatically correct and meaningful and align with the initial query or prompt so in other words you kind of separate concern you're dealing with the generative side so it takes the raw data from the retrieval model and give it more of a narrative structure too so then it kind of produce you know the content to be more digestible in that case um so essentially too over here is in the rank framework is that the generative models will give you it's like the final piece of the puzzle and gives you the textual output that you can interact with okay so now then real quick uh introduction to lstream so lstream is an again an open- Source library that my company uh data Stacks is uh working on our one of our teams are working on so this is just a diagram describing it too as you can see um lstream too essentially relies heavily on a event streaming type of model so if you're familiar some of you may be working with event streaming is basically you know um it it's like well I guess if I can say you know a platform that maybe a lot of you are familiar with would be Kafka for example so you basically can have data streaming in real time asynchronous data kind of streaming so that's the idea is that there are also connectors that we work with as well and we also work with llms all of the open AI uh vertex AI hugging phase all of these and also then it works also with Vector stores so not just with Cassandra or data stack our Astra DB but also pine cone and milish to these are other choices of vector database um again messaging is Kafka you can have it kind of plug in as a source of data feeding into your your model processing it and also PSAR as well and then also then it works with Lang chain like the libraries I mentioned earlier that works with um that you can build out templates or different types of things that you need to interact with llm llama index as well and all of these things and connectors too to other databases right behind the scenes there is snowflake or mongod DB all these things so it's a very flexible kind of library in there so over here too is basically works with Kafka and also too wanting to describe to you a little bit in a hurry is that it's composable too the model it uses composable agents too so in a pot you can have sources and sources can be kind of you you build like hefa uh Pulsar of the sources built in and you can build different agents that are processor that process all the data that comes into and then when you're kind of after processing you output all of your data to a sync like that so these are just list out all of the uh different uh connector connections it has so it's very flexible so also too the the benefits of using a library like Lan stream is that is declarative style and also low code too so you basically can capture all of your uh coding needs essentially in configuration form in yo file as well so so this is just an example of a low code yammo file example and here too is an example application for example a rack chatbot too you can have different pipelines that essentially do different things at the same time that feed into this chatbot interaction too so I realize I really don't have time so just want to really quickly kind of describe to with Lang chain what the problem it can help you with right so a typical pipeline you try to kind of bring data to a vector database for example you want it to be able to read data from a source so the source can be an s3b buet a website or Kafka topic or PSAR topic and then process the data ex and then extract the data to vectorize it and then you do need to also split the text into chunks of a given size and then compute these vector and beddings again you need to convert all your data and then you know and and chunk it everything and then write these chunks to the vector database and then clean up obsolute data from the vector database as you can see there are lots of steps that are involved in here so this is like a typical pipeline you know that brings data and then this is in yo file you can describe everything in a yamamo file too you don't need as much of a you know kind of coding because it's already there so just a quick kind of diagram describing it and this is what it is a rack AI applications and then you know you have you have you have this look in here so and I realized really don't have time so I'll really quick kind of like uh get to the demo part just wanted to kind of uh show you quickly too um let's see here uh where did to go okay so if you go to this uh site we have lstream do a I think langst stream. um or yeah anyway I have it on my slides too so I'll show it to you so basically too this is a quick start if you want to put your hands into working with this langst stream um basically you you know this example too is is so quick that it really doesn't take you too much time and you can you first have to install the Lang stream and then also have make sure you have open AI access key and whatever key you set up in your account and then essentially too you also want need to have darker running and essentially then bring this up too and then you can then interact with it so it's just a real real quick one so let me uh quickly kind of show it to you I already have this although I was having some issue earlier I think you can see that okay so all of the things I have already installed lstream I have my Docker up and all of these things so I just need to run lstream doer run test and that's what uh is required of this particular example and um and it will uh you know run and then when it's ready we also have a a simple HTML page that you can interact with the chat bot too so yeah let me see I like again I was having some issue with it I don't know it's it's always like that when you're ready to do demo then something goes wrong but that but this is the idea like if it runs too in fact I did actually a podcast uh with my colleague uh we were doing this podcast on gen and and then I actually showed this and it was right before or actually on the day of Halloween so I was just saying that please compose me a Halloween poem it actually did and it was really fun too it's basically using a Christmas song and kind of put in pluck in all the Halloween um kind of characters in there it was kind of fun too but let's let's see if it works now but if not you know you can always go to this site download your Lang stream and and try this out yourself too um okay so this is still while this is still running um okay so um I hope it's making sense because it's kind of short okay here we go so it it comes up which is good and let me kind of increase the so basically in here too um it makes use of you know you're producing messages is using Kafka in this particular example so you can connect to it and then oh this a status is failed not sure what is going on that's what was happen happening to me uh let me try it one more time this is kind of crazy but the but you get the ideas basically you can uh connect to it um uh sometimes it might have issue not not sure what is going on but the the idea is that you can then quickly say stuff like uh yeah you know uh write write a uh write a uh poem or something uh I don't know if it works but I think the the uh actually the producer status is failed so I'm not expecting this to be working uh in fact okay well I'm so sorry about that and I think I'm times up too because my hose is standing up right there so yes okay well anyway uh let me kind of go back to it but you get the idea you can quickly basically install lstream and then get your open API uh key and then have darker up and then you run this command uh langst stream doer and then you should be able to follow the instruction you be be able to bring this up and maybe I'll do it in our Q&A session let let me try to fix it and see what's going on okay anyway let me go back to my slides um where where's my SL can't find it okay I think maybe this this might be where it is too sorry okay okay here it's actually in the same browser okay so um and thanks for your patience with me so hopefully like I said on a Q&A session I'll fix that so thanks for staying and I know I'm running off time but I want to share with you the resources in here uh this slide deck can be accessed here if you are interested um and again you know if you miss it don't worry because I will be sharing the slide Deck with the organizers here too um okay okay all right thank you yeah thank you and then these are some more some more access and also our estra database I didn't even get a chance to show you we' like to encourage you to sign up because we give out $25 per month as a free credit you can use it for you know to build a vector database I think we allow you to uh for to build on your project and just to give it a try um that's a vector database on Astra and also I have my twitch stream this is my twitch um I'm trying to resume it very soon and then with that thank you very much and this is my LinkedIn and my ex uh Twitter handle and also a Discord server too if you like to stay in touch so thank you very much and I appreciate you coming to my talk thanks oh