Devreal

LLM Avalanche: Drew Minkin: Dolly-LLaMa: Unlocking LLM on a Budget

Event: LLM Avalanche SF 2023

LLM Avalanche: Drew Minkin: Dolly-LLaMa: Unlocking LLM on a Budget

Recording: LLM Avalanche: Drew Minkin: Dolly-LLaMa: Unlocking LLM on a Budget

thank you let me explain to you a little bit of who contact is we are an exclusive databricks solution integrator we care about databricks first and any other Technologies second doesn't matter what you bring if it touches databricks we will help you with it so just know that we're not selling you a product although I will tell you having been a CTO before that we engineer all of our consultancy work to be reusable it's kind of like our secret sauce you know not I'm not saying that uh for the private Equity people that we're looking to be the next DBT but you know you never know so before I say anything else um for those of you who are fans of uh clerks uh I wasn't really even supposed to be here today so I'm just gonna have to give a shout out to the uh the crew that made everything that you're going to see happen these are the people that we have that made the um the work that we're doing possible so I just want to make sure that you're aware that uh it's not just me and um to back it up a little bit you know I'm sure everybody here has heard of Dolly right from databricks is dead brick convention nothing new there but uh you know we take a different approach as far as how to get a good bang for your buck on llm right and so we're gonna take you on our journey that we've been through and we're actually going to show you a little bit of code right maybe more code than you've seen because you know we're here to uh to convince you that what we did was hard enough that you need us right I mean we're a consultancy right so just be prepared for those of you that are product Engineers you probably want to get another drink because this is not going to help you any right so anyway so um yes this is a a dolly generated uh image um I happen to be a little bit of a Tibetan Buddhist so it was uh my word play that made this particular name so we're very proud at least I'm very proud of it so anyway let's keep going here all right so there's been a lot of talk about cost and llm is being a big problem right and so what we've done is we've tried to take an approach to be able to make your ultimate fine-tuning process be available to run on CPUs right so in other words we want to make sure that even though you may be having a little bit of an upfront cost to be able to do your your initial model that your total cost of ownership is going to be a lot less right so that's what we'll be talking about today so I'm gonna I don't have a lot of time I'm a little upset because I lost a scarf uh that was given to be blessed by the Dalai Lama so if somebody sees that I am an alchemist as a I had an alchemist quote up there for those who didn't know Newton was an alchemist he died of mercury poisoning by the way Family secret for hundreds of years anyway point being is that um this is our process here so um you know we used a combination of traditional modeling techniques things that use databricks automl right things that we can um you know we'll you'll be able to recognize some of the code around there as we go through it and ultimately we're kind of at the Finish Line where we're using some ggml to take our Laura quantizations and be able to convert them into CPU right so just you know full disclosure the numbers that we have here what we're expecting to see but like I said Laura ggml is not a whole lot of a big deal compared to all the other fine-tuning that we've done here so I'm going to walk through a little bit of our process that we went through and here is I think this is big enough oh wait I'm sorry let me um I gotta switch to uh to duplicate here there we go all right is that big enough for you not drinking in the background I see two heads nodding and somebody watched it look in their watch so I'll try to get this done as quickly as possible Right so anyway so you know we started with the usual csbs we built some some nice little widgets there and ultimately we built a topic analysis you know standard LDA type stuff to be able to help us understand the um you know to give us a few more features to make it a little bit less work on the um on the modeling there and then one of the cool things that we did also is uh synthetic data generation if you haven't done this before in your own training we strongly suggest that you consider it because it's something that's helped us a lot uh even when we have problems that are not considered llm right the yellow synthetic day generation on things like um sentiment analysis um and uh classifications on text been very helpful and uh then we also took the um the topic analysis that we had and we built an automl around it in order to be able to build things and this is again before we do you know the you know there there's been a lot of talk about you know the you know a lot of Open Source models like bakuna and things like that and we're not saying that we've got a better uh formula per se what we're saying is that you know these are mini models that we want to have available for uh people to use because they'll be secure you'll be able to be complete controlling all the data going into them and um ultimately you'll be able to have your own micro model that we can fine tune and uh one thing I want to share that I'm very proud of and again like I said the the um the line on the bar chart is is is when we finish the um the actual conversion through the ggml but um let's back again yeah good okay all right so for those of you business-minded folk that are here which I probably all left to go um try to seal deals and not hang out with uh the propeller heads here you can see that essentially these are essentially the um our estimations for trying to use modeling we have a slightly different set of numbers that we came up with than other vendors and um you know they're a little bit more conservative I think they're a little bit more realistic but um you can see here that you know with our we we tuned on the uh the Hella swag score and we were able to beat the databricks dolly and we were able to shrink a model from the original Wizard of Acuna from 29 gigs to six gigs using these techniques that we talked about and when we finished the ggml conversion we are expecting it to be uh a six gig model there so um one thing I want to let you know is that you know we have uh you know like I said we are a consultancy but we approach our business like we were a software company and so what we've done is that we've built uh Dalai Lama Studio that we used in two-week kickstarts to be able to have customers that don't feel like they can they have the staff or the um the budget to be able to take on these big hundred thousand dollar things that we show them that a lot of times in two weeks you can find a lot of value and then be able to remove a lot of uncertainty for what needs to be done for bigger work and so this is one example that you know is directly related to our um our LM llm practice but I also want to put a little bit of plug because you know not everybody is ready I mean maybe everybody at this conference but not everybody is at the analytic maturity for doing llm work and so uh one thing I want to point out as far as another piece of work that we've done is that we have like I said we are we're 100 databricks SI and we've we're a deep partner with Sigma computing and so this is an example and sigma is not necessary for the the bi stuff but it's our preferred uh partner for along those lines what you see in the lower right hand corner is uh sort of an ERD of how we've built on top of the data bricks workflow and parameter system which would be the red and the orange boxes there a system of projects and phases to be able to have a very robust very observable way of being able to do work and be able to have great instrumentation you know integration Great Expectations all sorts of great things and last just want to let you know that we are doing an llm ask me anything we'll be doing live demos of The Dalai Lama Studio coming July 5th and so we just wanted to let you know that uh if you wanted to take a snapshot of the QR code that you would be uh we'd welcome you there so thank you all very much take care [Applause]