Devreal

AI For Good: Fighting Health Insurance with AI

Event: AI by the Bay

AI For Good: Fighting Health Insurance with AI | Holden Karau, AI By the Bay 2025

Recording: AI For Good: Fighting Health Insurance with AI | Holden Karau, AI By the Bay 2025

I'm going to be talking about fighting health insurance with AI. Um, and there's a lot of reasons why I care about this. One of the reasons is my dog uh had his uh claim denied as well as I've had several of my own claims denied. Uh but I am very protective of my dog and and I love him greatly. So yeah. Um so few few content warnings. Um I'm going to talk about surgery, hospitals, and health insurance. Um some stuff about being trans and some stuff about broken bones

And that's probably not what you are here for in general. And if that stuff sort of squicks you out, that's totally fine. Don't worry. Feel free to leave. I will not be offended at all in the slightest. Um, so yeah, I'm Holden. My pronouns are she or her. I am on the Spark PMC

Um, and I've been at a whole bunch of different places. Uh, and currently Gloria Stellar sign day job has has changed. um not yet using their name, but if anyone is interested in a day job working on Spark, come and find me out in the hallway later. Um co-author of several books. Uh most of them are about Spark and machine learning. I think they're pretty awesome. If you have a corporate credit card, I highly encourage you to purchase several copies this holiday season. Uh especially the outofdate ones

Uh those ones make excellent gifts and dogs love sitting in the boxes that they come in. Uh you can find me on the normal social media places and I also do a lot of live streaming including some of the work on uh fight health insurance. I I live stream it not as much as as I intended to. Um it turns out that like when you're dealing with health insurance it comes kind of allconsuming and so I I didn't stream as much as I wanted to but it is it is something where if you want to join uh you can watch Python versus the man. Um so I I did mention this. I'm trans, I'm queer, I'm Canadian. I got citizenship this year, uh, which is fantastic. Um, it feels much better

And I'm also part of the broader like leather community. These things are not overly related, but I think it's important for those of us who are building AI to look at people uh, who we're working with and say like, "Oh, wow. My AI startup is all people from the class of 2023 Stanford. Maybe we need to get someone from the class of 2022." um or more seriously, you know, get people from different backgrounds. Uh otherwise, we're just going to build sort of the same things that we've built but faster with AI. And I think having more diverse backgrounds present is especially important uh for those of us who are building AI tools. Um so the core problem is health insurance in America is terrible. Um and allegedly they're using AI to deny claims

that that has not been proven in a court of law yet, but I would be incredibly surprised if these allegations were not false. Sorry, were not true. Um, and so I think they probably are, but I'm not a lawyer, and my lawyer tells me not to say that they are using it until it's proven in a court of law. Um, so the other thing is that appealing denials is hard. It's really easy for a health insurance company to deny a claim, right? All they have to do is send you a letter that says like, "No, we don't think it's medically necessary." And then the amount of work that you have to do to appeal the claim is sort of asymmetric. On top of that, uh they have a lot more money than we do or at least than I do. I don't know if you have like Anthem Blue Cross money, come find me out in the hallway later. I would like some of that money

Uh but so there's this power asymmetry that makes this system especially terrible. Um and so my personal motivation, right, I got hit by a car in 2019. Don't recommend it. I got a lot of health insurance bills. Um, and then I'm also trans and I got a lot of insurance bills from existing as a trans person in America. And then also, as I mentioned, my dog, uh, who is perfect in every way, had his anesthesia claims denied. Uh, which is a pretty common thing for humans to get denied as well. Uh, because anesthesiologists are frequently out of network

Uh, not the same thing with dogs. They just don't want to pay for it. They don't even have the out of network excuse. Uh, they're just like, "No, we don't think your dog needs to be under anesthesia to get his teeth cleaned." And it's like, no, he's he's of like 10 pound Shelty. Uh, he doesn't understand a root canal. He's going to bite them. Um, and so that was a bit of back and forth there. Um, so I want to be clear here that AI can probably only do so much, right? Fundamentally, America is kind of broken in a lot of ways

And this is one of the ways that America is a little bit broken. We have a social problem. The social problem is people need healthare. and we've decided to make profit from that. And we've created this kind of Byzantine system where it ties this to your employment or to the whims of the government of the particular moment as to whether or not they're going to like claw back all of the subsidies that they gave you. And fundamentally, we probably need a different system, but we're going to do the flex tape solution, right? We have a broken system and we're going to try and work within that system to make it suck just a little bit less, right? We're not going to fix it, right? The there's still a giant hole in the side of the ship that is American healthcare, but we're going to make it leak a little less fast. Great success. Okay

So, how do we do this, right? Like, how do we potentially even begin to solve the problem that is American healthcare with AI? Uh, so we use machine learning. Yay, we're at AI by the Bay. You know, this shouldn't be all that surprising. Um, generally speaking, what we have is we've got this problem where we have all these inbound denials and people don't know how to appeal them. Um, part of that is like no one really sits you down and teaches you how to appeal a health insurance claim just like no one sits you down and teaches you how to file your taxes. America's fantastic. Um, and I can now make these joke is not quite the right word these statements because I'm a citizen. I feel very great about this

Um, and so that's cool. But if we build an ML model to appeal health insurance denials on its own and then just put it up on hugging face, it turns out that most of America doesn't know how to take a hugging face model and run it, right? Like that's just going to switch the barrier to people that are able to go out, buy an RTX 5090 and then download a model from Hugging Face, install PyTorch, have all that kind of fun and run it, right? So we we need to build a lot of other things as well as the machine learning model to actually make it accessible to people. Um and the final thing here is like PHI is really really important, right? So personal health information um it's it can be incredibly sensitive, right? Like I'm incredibly open about the fact that I'm trans, but there are other people for whom that could be a matter of life or death. Uh there are people who like may have diseases with certain stigmas they don't want the whole world to know. So, we need to find ways to make it so that we don't accidentally disclose their personal information. Um, and there's a whole bunch of tooling sort of around this. Uh, and we can do sort of a imperfect first pass, but so we're going to build some tools to help scrub out PII from data that we're receiving. And to be clear, I'm not saying that this is the perfect way to fix American healthcare, right? Uh, not even the only way to fix it with machine learning

Uh there are other people trying different approaches to fix healthcare with machine learning. I'm sure they're all great. Uh this is just the way that we've been working on uh right and then we have a stack of servers. At the end of the day, everything is a stack of servers and a very sketchy router which caused no end of amount of pain on Saturday night which I will in fact get to my my co-founder in the audience. So we can we can talk about that a little bit. Um so what do we need? How do we how do we go about creating this? So, the first thing that we need is we need lots of training data. It's like that scene in the Matrix where Neo is like, I need guns, lots of guns. Except that's not included here because that's copyrighted

Um, and unlike certain other companies, I do my best to try and respect copyright. Um, so we need lots of training data. So, the problem though is like where are we going to get this from? That's that's that's the big problem because I have my health insurance and there's a fair number of them, but 20 does not a training set make. Uh, we also need evaluation data. So ideally we need some way to tell are we doing good? Um and then we need a bunch of computers and we need a base model because we're probably not going to train a foundational model on just health insurance denials. Um the foundational model that would come out of that would be terrifying and also probably not very good and then some software to stitch it all together. So how do we get our training data right? Um, so first option is we go to the health insurance companies and we say like, "Hey, we're trying to make it so that you don't make as much money. I would love it if you gave me all of your data to help try and make the world a better place where your shareholders will earn less profit." Um, they did not in fact respond to my calls

Uh, not overly surprising. Um, okay. doctor's offices, the incentives are aligned, but understandably there are a whole bunch of laws about them sharing your medical records, and we want those laws, right? Like, we want that to be difficult. That is a thing that is important. Um, the other option is we can go on on Reddit and just be like, "Hey, what's up?" And it turns out that there are a large number of people who are so frustrated, so fed up with American healthcare that they're just like, "I give up. Dear internet, please anyone help me." Um, and that's cool. there is what's called sample bias and thankfully there's not like hundreds of thousands of examples of those with all of the data we need but there is some data there. Um and then we can go ahead we can look at the appeals that were suggested and we can write better versions of it ourselves

Uh and so that's cool but this this is still not giving us a full training data set. So the the spoiler is essentially we go to these things called independent medical review boards. And so each state uh because this is America and we like giving states power uh to do things has the ability to set up independent medical review boards. They'll call them by different names. Uh for example, if you search for independent medical review in Washington state, you'll find something about workers comp, not the HHS version, but the California HHS has all of their decisions published. Um similarly, New York has their data published. Washington state does have it published. It's under a different name, but this is fantastic

And it is it is already somewhat anonymized, right? So, this is fantastic. We don't have to worry about accidentally leaking people's personal information by consuming these records because they've already been anonymized and they're already being shared or deidentified. Sorry. Anonymized has a slightly different meaning. Um, and in fact, there's there's people who have made tools for exploring this. So, you can go and you can see how um borderline evil the insurance companies really are. Uh the the appeal explorer was is a new tool and it's pretty cool. It's only on New York data, but you can go and you can look and you can see like, wow, insurance companies really don't want to pay for preventative cancer screening

What the hell? Um it's terrible. Uh okay, right? And so then you know what do these records look like? So these records are fantastic, right? But they're still not quite what we want, right? What we want is we want denials and appeals. Uh instead we've got C. Um, so right, we want A and B, but we've got C, which is the product of A and B. So, okay, we've got something that contains most of the information we want, but not in the format that we want. But this is fantastic because this is actually one of the things that large language models are pretty good at, right? We can go we can take C, which is the product of A and B, and we can ask other large language models to produce what A and B might have been. And this actually works pretty well, right? Um, super simple prompt. We we played with a whole bunch of different prompts

I don't think this is the current prompt that we're using, but it's it's a very simple idea where we take the findings, right? This is what the final appeal found and it references what the insurance company did and why the insurance company was wrong to do it. So, we can go ahead and we can go back and we can figure out what an appeal would have looked like and what the denial would have looked like from this. And by we, I mean the computers. And that's fantastic because the computer can generate thousands of these and I can generate like 10. So that's fantastic. So we put this together with our traditional big data except it turns out we don't actually need Spark. Um which as someone who works on Spark is a little depressing. Uh we need for loops

Uh four loops are pretty good when you've got thousands of records. Even tens of thousands of records, yeah, for loop is still probably going to be pretty great. We don't we don't need distributed computing at this point. Um, oh, and here we've got my very sad Vespa. Uh, this is what happens when a Vespa and a BMW say hello to each other. Uh, the Vespa, the Vespa looks like this and is no longer ridable. And the BMW is repaired without incident. Um, the rider then gets to go to the DMV and get a very special placard

Uh, and and so for a while I got to park in all of the fancy parking spots. still not worth it, but okay. Okay. Uh some benefits. Some benefits. You know, you get these these stylish stylish devices. Um and so very simply, right, we're we're just going through, we're iterating through, we're calling out to the model and being like, "Hey, what's up? Did this work? Like, go ahead, generate an appeal and a denial for me." And then of course we've got like these weird formatted targeted files because we want to try generating with different models and then we're going to look at them and so that's why we've got these weird format strings for for what the what the output is. And then finally right um everyone knows LLM or most people are probably aware LLMs are not always amazing and sometimes will just produce garbage

So then of course we do what everyone does when encountered with a problem is we add regular expressions and now we have an uncountable number of problems instead of one problem. Um so essentially we check for the results coming back from the LLMs and say like hey does this look at all reasonable right like did it just go off and like hallucinate a whole bunch of random stuff. Uh and that's that's a very manual process where you like review a bunch of results from the LLMs and then you say like okay these are the places where I see it misbehaving. Let's try and write some rules for this misbehavior. Cool. So now we've got a data set though. That's fantastic, right? It did not cost hundreds of thousands of dollars. It cost hundreds of thousands of like I don't know GPU seconds

Um which thankfully is is still much cheaper. Uh, and it did cost us money to generate, but it turns out if you're nice enough, uh, various cloud companies will give you free credits to use their GPUs. Uh, shout out to Microsoft being the one who actually had GPUs that we could use. Um, they were they were very nice. They let us use their GPUs. I don't know if they want that shout out though. So like maybe uh I don't know how closely they wanted to be associated with this. I should have asked for

Oh well. Okay. So we've generated some data. Now we have to filter it out because the robots are not perfect and the other thing is the licensing is important. Um now I am not a lawyer and there is some question as to are machineenerated works copyrightable and I think the answer is generally no with an asterisk but I'm not a lawyer and you should always follow the terms of service of whichever model you're using. And in fact, that means that before you go ahead and run this handy dandy for loop, uh you should go ahead and check and make sure that in your list of models is only models where you're allowed to use the output. All kinds of fun. Thankfully, there are many many models where you are allowed to use the output to train another model

Fantastic. Um so like don't use OpenAI. They'll get very mad and probably call you bad names in a press release. So [snorts] this is cool. We've got a data set. How do we go ahead and actually turn this into something that's going to work? So now we need to pick a model to fine-tune um how are we going to make that choice? So the first thing is is it going to fit in memory right? So if you have an infinite amount amount of money like if you're a hyperscaler like Google or Microsoft not not as much of a concern. On the other hand, if you're like a, you know, punk startup in the mission, you're maybe looking at like, okay, cool. I've got two 3090s each with 24 gigs of VRAM

Let's make sure our model is going to fit in under 40 gigs of VRAM, right? Like I I really I want I want this to be able to work. Other thing is we also have to check the license and we also want to pick a good base model. Ideally, if you can find a base model that is kind of similar to the task that you're trying to accomplish, that's probably going to be even better. But it's okay, right? Like there's there's no other generic health insurance appeal model. There's no current like h Did anyone here watch Yes Minister growing up? There's zero people. Okay. It's a fantastic British comedy. Oh, we've got one person

Okay. So for the one person in the back, there is no Sir Humphrey Applebee LLM out there yet, but I want there to be. Okay. And they are putting their like this. Uh which is how you know it was a very good joke and you should all laugh. Uh because otherwise I will feel sad. Okay. [snorts] Um anyways, so we we'll we'll pick a model

There's nothing that's like a perfect fit to this, which is why we're fine-tuning something anyways. If there was a perfect fit, we probably wouldn't need to fine-tune it. Um, so how do we do the fine-tuning? Right? So we already we we've got our data. We need to structure it in the correct way which is mostly just annoying JSON and then piles of shell scripts because the world is fantastic and still runs on piles of shell scripts and Python scripts. And then oxalottle simplifies a lot of the shell scripts that I wrote initially. Um, and if you're looking to fine-tune a model, I highly recommend them. Uh, it works surprisingly well, especially compared to almost all of the other tools that I tried, which would normally work for like a week and then it turned out that they wouldn't have any of their dependencies pinned and something would change and then then everything would just go to hell and they'd all break. Uh, and Oxodle pinned at least some of their dependencies, so it kind of worked almost repeatably

Um, my co-founder can attest to almost being the the operative word. So, okay, how did we make our most recent model? So, we, you know, fine-tuned this model on our synthetic output. Um, and we used Lambda Labs to do some of the fine-tuning uh for the first model and as I mentioned more recently, we used the Microsoft Azer one and then we used Oxalottle. Very cute little logo. Super happy here. Um so right now we've got a Gemma derived model. Before that we were using a Mistral derived model. Um and we are in the process of deploying a medge gemma derived model and I'll talk about that experience later

Yeah, I still have time. Oh maybe not a lot. Okay. Uh TLDDR bunch of shell scripts. If you're the kind of person that likes shell scripts, great news. Uh they're in our health insurance LLM repo which I linked to at the end. You know uh configuration. Everyone loves YAML apparently

So it's just YAML to sort of configure the base model super easily. You know, you point this to the data set that you've generated and then you go ahead and you run it and you have some magic tokens. And these will all depend and have to be set depending on which specific base model you are fine-tuning. And here I am sad uh with a lot of stuffed animals, which is weird because normally I'm happy when I have a lot of stuffed animals, but I think it's that I can't move my arms to drink the coffee. So that was a very sad day. Okay. So fine-tuning a model costs about $112.32 give or take. Uh it costs more if you do it on uh one of the big clouds, but you can get credits

So fantastic. Uh now we need to serve our model. So where are we going to serve our model? So we're going to use a rack in Fremont. Does anyone else have a rack in Hurricane Electric in Fremont? One person. Do you need BGP transit? We'll talk afterwards. Okay. Side business, don't worry. Nothing to see here

Fantastic. I, you know, in in between my like day job appealing health insurance denials, I also run a small ISP. Uh because who needs free time, right? Uh so things that we learned, RTX3090s do not fit nicely inside of 2U servers. Um market segmentation in this case by physically having the card be too big. Uh fantastic. And it uses power and costs money but not not that much. Uh fantastic FCX amazing people don't worry about it. TLDDR we can configure our model serving using VLLM

There are of course different tools that you can use. It's the one that we used. Uh we did find that nothing worked on any of the ARM machines we had. So we used AMD 64 machines instead. Cool. Super fun. Enforce eager. Super important

Otherwise, it'll just fail randomly uh later. And so, I would rather things fail at the start when we try and run them. And this costs about $74946. Uh eminently affordable. Unfortunately, the racks are a little bit more expensive. I have some sketchy plans. We'll we'll talk about those later. Um but Medma does demand more me uh more VRAMm than you can fit uh in a 3090, but it does fit nicely in a 5090

and Microsenter had one on sale. Uh so it was not $5,000. I think it was $4,000. It was fantastic though. We got a bigger computer. Everything's amazing. So uh then we have to build a front end because as previously mentioned, people don't want to call like use curl to appeal their health insurance denials and submit a JSON blob. They want click button receive healthcare

Um, and internet access super overengineered for the one person in the back. We'll talk later about this. Uh, now I was going to do a demo, but we have seven minutes and I do want to save a little bit of time for questions, but trust me when I say it totally works and there are actually people in this room who have used it who are not me who can attest to the fact that it has worked. Um, so actually we'll do the demo if there are no questions. Um, the other the other thing is going from lab to production continues to be kind of rough. Uh, so this is Disneyland, which is the lab. Super happy. Food's delicious

And then we've got Burning Man over here. That's production. Uh, everything's on fire, literally, because you brought a bunch of propane and decided to light it on fire. Um, and so we had an outage on Saturday night because we were upgrading something. Everything looked great and then everything fell over and cascaded and we ended up having uh the load balancer software that we were using was acquired by another company which decided to pull all of the images from DockerHub and so everything just broke instead of restarting gracefully. And I was very sad and it took several hours to make it all come back but it did. Um, and we've got a new model coming soon. It's going to be so cool

And if you are the kind of person that's like, "Fuck the man. I think people deserve healthcare." Uh, PRs are welcome. So, you too cannot have a life and spend your free time on Saturday nights debugging random or contributing to giant Python uh projects. Fantastic. Also, if you work for a cloud hyperscaler and can throw us money or credits, uh, give me a call. I like free things. Um, and yeah, okay, five minutes. Here are a whole bunch of links

Uh, you can check these all out. Um, I think you should totally do that. Pull requests are welcome. Really actually welcome. And if you want to like talk about where you can help, come and find me. Uh, I'd love to talk about it. And now I think it is going to be time for questions or a demo. Who wants to ask a question? Okay, we've got no hands for questions

So, we are going to do a demo and uh if people change their mind partway through this demo, not my fault. Okay. So, here we can see why back-end developers should not build front-end applications. Um so, we're going to go ahead and we're going to generate a health insurance appeal with AI. And so, here we're going to go we're going to say my name is Holden. Actually, we're going to say my name is Timbit Caro. My email is Timbitimbit.com and my street address is 123 Fake Street. And the interesting thing about these is like we don't actually transmit the name, email or street address to the server

Those are all actually client side uh stored. So when we go dear Timbit uh your anesthesia was denied as not medically necessary uh during your root canal. Uh sincerely Dr. Evil Evil Insurance Corp. Definitely not Anthem. Okay, fantastic. Oh, and then the, you know, ooh, regular expressions. Um, and then we obviously will read all of these policies super carefully before we click to agree as everyone does

Um, and then we'll click these two. Yeah. Okay. Cool. And then option relevant health history. I am my dog. Um let's see here. Uh [snorts] cavity

Cool. If we had our plan documents, we could upload them here. It turns out almost no one has their plan documents. And in fact, it's incredibly difficult to get them for many plans. Um, most ins employer sponsored health insurance plans are required to provide a full copy of your plan documentation within 30 days of written request. I have yet to find one that actually does. Um, it's fantastic. I'm very sad

Okay, cool. Uh, here we go. We're doing some magic ML in the background, but uh, we're just going to go ahead and we're going to hit skip. Um, okay, cool. Unfortunately, it picked Anthem. Who would have known that evil insurance corp wasn't in its database? I don't know why it would pick Anthem anyways. Okay. Uh, and we're going to say that we have an employer plan

Timbit works for the government. That that seems that seems reasonable. Um, and we are not robots. Okay. I did not successfully extract this. Uh, anesthesia. Cool. Super happy fun times

Okay. I hope it works. Cool. And now it asks a bunch of questions. And these questions are generated by the model where it's like, hey, uh, we don't really know what's going on. Um, and so yes, um, uh, let's go with no. No, no historical allergic reactions to anesthesia. And yes, they are a dog

Um, they're a dog and would bite the dentist. Okay, cool. And now we go ahead and we generate our health insurance appeals with a very cute picture of my dog who I love. Uh, and then this might take a minute though, so I really hope it works. Uh, sometimes it doesn't and then you have to hit refresh, but okay. Uh, we've got like 1 minute, so I'm going to call it a day if it doesn't finish really soon because this this step can take a few minutes, unfortunately. Uh, but I assure you it does actually work. It just it just does take a bit of time sometimes

Uh, okay. Uh, it did not work today, but it normally does work, and I will figure out why it did not work later. Uh this is how you know the demo was definitely not faked though. So fantastic. Uh we'll we'll figure out why it it did not want to generate the appeal properly uh right now. It could have been that I I used the dog example. Uh but more likely it's that like the server is having a sad time. So we're going to call it a day

I'm going to be around I think I'm on a panel at 4 o'clock about the future of the AI stack. So, you should definitely come and say hello to me there, too. Uh, thank you all for listening. Sorry, the demo crashed. >> Thank you. Yes, thank you. Well, since it's lunchtime, if you want to ask a question before we, you know, go to lunch, that's fine. Anybody feel like wanting to ask a question? >> Okay

>> Uh, are uh are people using it? >> Yeah. Uh that's that's we've got we've generated a few thousand appeals so far. Um so people are using it. I I wish we were generating more appeals. Uh but yeah. Yeah. >> Uh so uh do you have like any metrics around how many of those are successful or are those still ongoing? >> Yeah, that's a good question. So because of how we do it, uh we generate the appeal and then we give it to the patient to submit the appeal themselves

Um, and because we don't store their contact information unless they opt in, uh, there is a checkbox where they can choose to let us contact them afterwards. Uh, we only have like sort of survey information. Um, but of the people who respond like yeah, there there seems to be a fairly high success rate, but there is a huge amount of sample bias there where it's like they both have to select that it's okay for us to contact them and then answer a random email like several months later asking them how it went. Uh so I I would be cautious about drawing any strong inference from that. >> Thank you. Thank >> you. >> Yeah. Thank you all

So right now it's lunch and then 1:30 we'll be back with a keynote right after lunch. So yeah. All right. Thank you. Thank you all for coming and thank you to uh Holden for the presentation. [applause]