data.bythebay.io: Eric Williams, What Healthcare Can Learn from Netflix
Recording: data.bythebay.io: Eric Williams, What Healthcare Can Learn from Netflix
Hello. Hi. Thanks for having me. Um, like you said, my name is Eric Williams, director of data science at Omada Health. Uh, my somewhat sensationalistic headline here, what healthcare can learn from Netflix. really want to talk about what how to build personalization and optimization into in our context preventative care but really any intervention any interaction you have with your user consumer etc. Here's the outline first give you context about then talking about kind of how do you build a data science culture because I feel like to really get at personalized anything organization team you need data science kind of boiled in from the ground up. uh then a couple specific examples about how you actually go about datadriven personalization
So first AMA this our mission statement we inspire and enable people everywhere to live free of chronic disease. chronic disease being the obvious problem. Um, you know, classic iceberg slide, uh, nearly 30 million Americans with type two diabetes, but the, uh, more threatening impact is that almost 90 million Americans on the clinical brink, pre-diabetic. Lots of these people don't know it, and it's estimated about 5 to 10% of these are going to convert to type two per year. Uh, the solution, I think it's kind of known. I think if I asked each one of you, you know, the solution to some degree. It's about lifestyle change. This is uh data from the diabetes prevention program study which is a clinical trial that showed really validated that health coach le intensive behavioral uh counseling should be standard of care for those at risk of weight related chronic disease
Uh so it showed the lifestyle arm intervention arm outperformed the the control arm by about almost 60% in reducing the risk of developing type two diabetes after three years. It also outperformed the standard prescription drug approach called metformin. So kind of begs the question if we know the problem the solution why is there this problem? Uh I think everyone also kind of knows this behavior change is hard. Um I think everyone kind of resonates with this on a personal level. There's a lot of dimensions to it especially if we're talking about kind of the digital health aspect of it. Behavior change in a digital context. There's obviously an educational component. There's tracking
This is an obvious one. Uh really popular right now. Calorie counting is often touted as a as a solution. Health coaching uh in the health care space. This is kind of hot right now. Social support, especially when you think about the difference between social support in person versus online. What does that mean? Where's the efficacy there? And kind of the argument that I'm making here is that you need each all of these together kind of in in a symphony. Um, and that kind of leads to the the company or the OMAD program
Let's see if this works. There we go. Program. So, we're are an online diabetes prevention program. So, people coming into our program, they get partnered with a health coach, a small group of about 20 people. So, imagine kind of a mini Facebook type environment for social support. We mail them a physical scale that lands in their bedroom, bathroom, uh, kitchen that we read the their longitudinal weight data from. It's obviously available online on the on the mobile app or the web, iPad
Uh, two business related points. One, we're B2B, so we sell the large self-insured employers or or health plans that make it available to their their uh, employees or their members. So, no one on our program actually pays for it. Second point is we're outcomesbased pricing. So we don't actually make money unless the participants on the program actually lose weight and thus actually reduce their risk of uh chronic disease. Just to uh support this idea of needing the full symphony here are some of our weight loss outcomes. Y-axis is percent weight loss at four and 12 months. The blue and the green curves are our commercial and our published results
So we do publish and peerreview journals. The red line is uh results from inperson versions of us. So inerson face-to-face diabetes prevention program they have about two and a half% weight loss after a year where ours averages about 5%. And the orange is kind of the employer wellness program version of this. So back to the concept uh the whole symphony what I've added here is data science kind of acting as the conductor orchestrating the right intervention at the right time for the right person. And this is kind of the core of preventative uh personalized preventative medicine. You also hear the term precision around this although precision often invokes the uh genomics. If data science adomada had a mission statement, it would be about using analytics, machine learning and experimentation to really parameterize what works and what doesn't for behavior change
So real quick side point, you know, these are standard data science toolkit um are generally used to optimize interactions with your participants or your users or your patients to optimize click-through rates, funnels, ad conversions. We're using the same toolkit, but using it to optimize the reading coming off the scales. So weight loss. So it's data driven weight loss, data driven chron chronic disease reduction. And I've said that enough times. Right intervention, the right time, the right patient. So, a little bit about our data. So, this is a snapshot of a subsection of data from one participant out of 60,000
It's kind of funny depending on the audience. 60,000 is either either a lot if I'm talking to a healthcare audience or it's a very little talking to a more tech audience. Uh, but in behavioral science, this is a massive coherent data set. So, what you're looking at is one one participant time in the programs on the x- axis. The right hand side y- axis is weight loss which corresponds to that diagonal fitted line which means that this participant after about 17 weeks lost about 13% of their body weight. Everything else is kind of interactive data that we collect. Each dot it corresponds to one of the the um interactions on the left hand side y-axis. So for example, private messages between health uh sorry priv private messages between participants and health coaches
a lot of emotionally charged text data that we we like to dive into about problems with behavior change, struggles, physical activity tracking. So, I'm going to talk a bit about this. Just like your jawbone or your Fitbit, we collect steps, key component in behavior change. The social dynamic, again, a lot of emotionally charged text between participants, meal tracking, um over six million meals, uh with healthiness, portion size, etc. And I just picked out a few there I wanted to highlight. So it's kind of a picture of the data the uh context I want to talk about data science culture in healthcare in particular it's kind of a new new concept uh imaginary scenario you're the first data scientist at your company you know you have amazing data a lot of potential but you're your your company doesn't really have a datadriven product development cycle or no data science team where do you start uh these are just some of the things that I tried they each worked to to a degree One of these kind of jumped out had a oversized ROI. Believe it or not, it was the data blog and I actually wanted to go through this a little bit because it does give good insight into our data. Called it plot of the week
This was just an internal email weekly blog buzzfeedy type catchy stories about our data really just to expose the company to potentials of our data and then mobilize support and excitement. So I want to very quickly blow through three examples of these. The first one remember their scales arrive at the doorstep of our participants and really become part of the family. And we capture the data of anything that steps on that scale. Each scale is associated with one account, but we very often see things like this. Days are on the x- axis, weight is on the y- axis, and this is just a family of two losing weight together. It's pretty common to see. It's nice to see, too
Here's a family of three of some weight fluctuation, but you also see a growing child over about a year and a half. on the bottom there. You can also see some of our noise. Here's another family of three that looks like they've had the neighborhood over for a party on uh day 400 to jump on the scale. Uh here's a family of five and it looks like the uh the main participant forbade his families to step on the scale for the first four months of the program and then everyone started using it. Uh and then there's weird stuff like this which I just don't know. Looks like they're using the skill in the wrong dimension or something. Um, my hypothesis is that this is some sort of school, maybe a maybe a Sunday school or something
You have heavier adults, lighter children, and some periodicity to it. Okay, I'm going to try to speed up a little bit. Next one, Minds Over Matter. This is based on this book called Mindset: New Psychology of Success. Does anyone know about this book? Yeah, one or two. Fascinating. Strongly recommend it. Um take-home point is people come in two flavors
Fixed versus growth. Fixed mindsets think their internal characteristics are part of them. They can't really change growth. You can change these things through practice exercise. So a program like ours is primed to measure this. We asked ourselves can we identify our participants mindsets in the program and then look at how they do. So a fixed mindset actually associates negative trend negative tendencies or negative traits with u the present state of being. It is hard for me to lose weight while a growth often refers to these traits as part of the past
I used to struggle with my weight. So we have all that text. We can create a simple mindset sentiment score. We did that. Here's just two participants. The upper plot, both plots are time on the x-axis. The upper plot is actually um two participants. Each dot is a private message to their health coach
The y-axis is percent of that message framed in the past. So if we if we use that as a proxy, the h the more it's framed in the past, the more growth. Same participants on the bottom with weight loss. We're just seeing somewhat cherrypicked example of uh the growth mindset losing more weight. Just a proof of concept. Definitely not a uh a thorough analysis, but again, it was a hearts and minds campaign for these blog posts. Last example really quick starts with a quote. I only weigh myself immediately after I wake up
I'm lighter in the morning. I don't care what you say. I'm lighter in the morning. Uh this was from my wife. When I started working with weight loss data, I thought I might as well at least Google it. Turns out every night you lose more than a pound while you're asleep for the oddest reason. Does anyone know that reason? Just shout it out. I might have heard it
It's uh you actually you breathe it out. You we lose about a pound each night from breathing out carbon and water vapor. Um sweat's part of that too, but the vast majority is our breath. So our data set's actually primed to measure that. Look for people that weighed in late at night again early in the morning. Look at the distribution. Sure enough, we saw about 1.3 pounds of weight loss overnight. If you stratify that by time, duration or proxy for amount of sleep which you see that difference increase to about eight hours where it turns around at 9 hours or maybe people are coming to the scale rehydrated with breakfast
Okay. So now I actually want to get into personalization. The first step to personalization is really uh I believe it's experimentation. You have to be running experiments. randomized control trials I think could go on a whole different talk about how digital uh AB tests digital health can really revolutionize the old uh paradigm of randomized control trials but I guess the bottom line here is if we can ask fundamental questions about behavior change and run these 60,000 participants through our program through this experimentation we can really set a new velocity for behavioral science discovery um and massive personalization comes comes along with that kind of the two classic ways or at least two ways I want to talk about first is a subgroup analysis run an experiment and then you ask questions like does this affect older people differently than younger people males versus females uh obviously this is guided by heruristics I think sometimes this can be limiting and lead to biases but it's very powerful too one example of this in our program physical activity is obviously a large part of healthy lifestyle we focus on physical activity for a good a good portion we actually mail our participants a pedometer so we start collecting steps they can also link Fitbit job own etc. Uh we also combine that with educational support from the health coach and materials. And then we set daily step goals for participants. And this is where the data driven part comes in
Um because like a jawbone, you get 10,000 steps a day. But we asked ourselves, is that reasonable? Uh we looked at the data. On the x- axis, this is our participants BMI. Y axis is historical steps tracked. And you see after a B above a BMI of about 35, probably for solid physiological reasons, their their daily steps that they walk start to drop. So maybe 10,000 steps per user isn't the best idea. Here it's it's too small to see. I get that uh there historical distributions of steps segmented by BMI horizontally, age vertically
And so we just came up with an algorithm for a new participant coming in to assign them a step goal. We look up their age and BMI and we look historically at the average steps tracked by people like them in that age and BMI. Add 20% to that and give it to them. We did this as an experiment. Half the people got that adaptive or I think I call it personalized. Half people got a static of 7500. X-axis on both these is program week. Y axis is steps tracked per week
Females on the left, males on the right, personalized in blue, static in green. Not quite conclusive yet, but we could be seeing hints that these personalization really uh works is more effective for males. This is also a age segment. These are these are the youngest participants we have 18 to 30. If we zoom out and look at older participants, 50 to 55 for example, that difference really disappears. So it's just an example of running experiment subgrouping and seeing where it's most effective for and then we can personalize after that. Final example of how you can use experimentation to get at personalization is something called up uplift modeling. Uplift modeling is a statistical model based on experimental data
So you need experiments to start. It predicts who the intervention is likely to be most impactful for. And the thing that I really like about it is you don't need heruristics. the the data will do the personalization for you. So as another example, nutritional awareness is al obviously also a large part of healthy lifestyle. Our program focuses on food tracking to raise nutritional awareness. The experiment was pretty simple. We wanted to understand the effect of a coach, a human being giving feedback on a person's food tracking
So two arms 5050. Uh both arms are tracking food. One arm that food is uh um surfaced to the health coach and that health coach gives them feedback kicks off a conversation about the specifics of the food they track. The other arm is tracking their food. They're just not receiving feedback. Kind of the more traditional food tracking. Pretty simple. Roll that out
I saw really amazing results. Um 10 to 15% increase in participants fraction participants who tracked a meal. was powered at over two months. 8 to 12% uh increase in those meals with healthiness tracked. And then the most important thing was we saw an 8% relative increase in weight loss at 16 weeks. It's a relative. So that means if average in the control group is 5% weight loss, intervention group was 5.4%. For for us and for diabetes risk reduction, this is really huge
So we know that this intervention works on everyone. Are there people that it works best for? And that's where uplift modeling comes in. Did it work? Uplift modeling is really for datadriven personalization. It rests on this assumption that your patients or users or customers or whatever come in one of these four quadrants. Sure things are people who are going to respond to your intervention whether or not you give it to them. So in our case, these are people that are going to keep tracking their meals whether or not the health coach gives them feedback. The lost causes are users that are going to not going to respond regardless of the intervention. In our case, these are people that are going to stop tracking meals whether or not the coach reaches out to them
The persuadables are only going to respond because of the intervention. So in our case, these are people that are only going to continue tracking meals because the health code is reaching out. And then the sleeping dogs, these are people that are less likely to respond because of the intervention. This was me when Zipar sent me a reminder about my subscription to Zipar that I forgot about and all that the the reminder did was remind me that I didn't want it so I unsubscribed. Maybe in our case, this is someone who tracked a you know their willpower broke down. They ate a Big Mac. They tracked it. They weren't proud of it
They knew it wasn't healthy. Our health coach reaches out to them about that Big Mac. They become ashamed and they stop tracking. The whole point of the Uplift model is to target the persuadables and remove the sleeping dogs. So we built a model on our our data. There's a lot of ways to do it. We chose a random force uplift model trained on the likelihood of participant continuing to track meals after receiving food feedback. Inputs to this model are participant demographics, program behaviors, the food they tracked, response to previous feedback, etc
And the results showed that using the that model to target the top 80% of people that are most likely to continue tracking based on what the model says the we increased our response 2%. And it means a five that's relative 2%. So from 5% to 7%. Um in other words the the bottom 20% that the model said don't target those are the sleeping dogs. Those are people that would have stopped tracking if they had been reached out to. Very powerful. you do need a lot of data um a lot of really clean experimental data perfect so that's that's basically what I want to talk about give you a little context about health or doing datadriven chronic disease prevention uh what my thoughts on kind of building a the importance of building a data science culture in your organization if you're going to do these kind of soup to nuts experiments to get at uh personalization because you have to start with the experiment follow through and then go to personalization and then finally some details on how we actually do this at Amada. Thank you
[Applause]