Devreal

Continuous Delivery Principles for Machi...

Event: Scale by the Bay

scale.bythebay.io: Rajesh Muppalla, Continuous Delivery Principles for Machine Learning

Recording: scale.bythebay.io: Rajesh Muppalla, Continuous Delivery Principles for Machine Learning

Hello everybody, my name is Rajesh Muppala. So today I'm going to talk about Contest delivery principles for machine learning. Or in other words, how do you solve the last mile problem of taking machine learning models to production? So let's get started. A few things about me. So I'm a co-founder at Indics, where I lead the catalog and the machine learning infrastructure team. And I'm from Chennai, and I'm actually flying all the way down from Chennai to here. So I thought it's a good time to actually talk a few fun facts about Chennai. So Chennai is about 200 miles to the east and north of Bangalore

is actually called the Silicon Valley of India. And a lot of people don't know that actually Chennai is north of Bangalore. And even I didn't know because I moved from Bangalore to Chennai. It's also the home of S. Ramarjam, a genius mathematician from 19th century. And also Sundar Pichai, CEO of Google. And we have three seasons, hot, hotter, and hottest. But I like it there

I've been there for six years now. This is my second time at Scale, by the way. I was here last year as a panelist on data pipelines. And there's an interesting back story here. And I met Alexei about three years back on a plane. And I showed some of the stuff that we were doing at Index. And he said, come talk. And then that's how I land up there

And I really like the community. I mean, last year I made a lot of friends. And I think this year it's not going to be any different. And previously I've spoken and written about microservice and lambda architecture, again, in the context of how we apply them at Index. And before Index, I was at ThoughtWorks, where I was a tech lead on an open source product called GoCD, which is a CI, a Cundice integration and Cundice delivery tool. So it looks like a few people have heard about it. A quick show of hands, how many of you have heard about Index? One, two, three. OK, I expected that

So let's talk about Index. So at Index, we believe there are six critical indexes for business. For people, you have Facebook, LinkedIn, and Twitter. For documents, you have Google and Microsoft. For businesses, you have DNB, WhoWars. With IoT devices coming up, you'll see an index of connected devices. Somebody will build it. And for places, you have Google Maps in here

And we want to be the index for products. So you have Walmart and Amazon, who have their own index of our catalog of products. So the issue there is that companies have to be in their ecosystem to use their catalog. I mean, at Index, we're trying to build a catalog that anybody can use. And you can look at us as the Google Maps of products. So if you realize, now, turn-by-turn direction on navigation is a solved problem, because Google Maps took the hard problem and solved it for the companies like Uber and Lyft and Ola. We want to do, I think there are companies like Uber that are out there who would benefit from somebody solving the problem of index of products. And that's what we are doing

And it's a pretty ambitious goal, if you think about it, being the catalog of all the products in the world. And we've made some dent at that. We have about 2.1 billion product offers already, which is one of the largest collection of products in the world. So how do we actually gather this data? So let's talk about the data pipeline at Index. We gather most of this data through crawling brand and retailer websites. So we have this process called seed crawl parse. Seed is discovering URLs. Passing is actually going and crawling them, and parsing is converting the structured HTML pages into product data

We also have some data that comes through brand and retailer feeds. And then we have this, what we call ML data pipeline, that merges this data. The first thing it does is dedoop, because we see a lot of duplicate pages, URs pointing to the same product on web pages. And it also dedoops between crawl and feed. The next step is classify. So we have our own taxonomy, because we need to classify a product to our taxonomy. So that's the classification engine there. We also need to extract attributes

For example, if there is a product called Apple iPhone 64 GB, so the fact that Apple is brand, iPhone is a product line, and 64 GB is storage, is something we extract in that. The next step is standardizing. The attributes might be called differently on different websites. So how do you standardize them? Then matching. The product might be sold at multiple websites, so making sure that they are integrated so that all analytics further can happen. And then we expose this as product data transmission service. So if companies want to use these APIs, they can go ahead and use it. We also make this available as a catalog through customizable feeds

And finally, we index this data, and companies can use our API to search and use this data. So who uses us? So we have retailers and marketplaces who use it. So if you see some of the companies here, they are traditional brick and mortar stores. And one of the challenges they are having is they have products sitting on the shelf, but they can't put it online to compete against Amazon. The reason being that their supply chain is not optimized for the digital world. I mean, you go to a store, you look at the product, you buy it, you have manuals, but they don't have digital information. That's where we come in, where we fill in the holes. Like for example, if this is the catalog we get from one of these retailers, we actually fill in the holes as well as we provide additional data

And this data is also SEO optimized so that they can put that product online. We also have ad display and exchange platforms using us. So again, similar use cases. So advertisers, retailers, and publishers, we do enrichment. So the clear goal to provide more personalized ads. Apart from this, there are a few interesting use cases that I would like to mention. So we have insurance companies that use our data to price the premium of the products, because we give them price information. We have companies that scan your receipts and tell you how much you spend in each category

So they take care of the OCR and they use our classification to figure out the category. And more interesting, I mean, there's an interesting use case that I want to talk about. So UPS and FedEx use us because it turns out that a lot of packages, the labels in packages get lost. So they want to figure out what's inside the package. So they build an app that uses our APIs to scan the product, UPCs or barcodes, and then figure out and catalog it. So earlier they used to use Google search for doing the same. So yeah, so you see that there are interesting use cases once you have product data coming up. This is the scale at index

I spoke about the 2.1 billion product URLs. We crawl close to 8 terabytes of HTML data, which correspond to about 30 million URLs per day. We have about 3,000 sites in our catalog. And we have about 7,000 categories that we support. These are leaf level categories. Okay, so this is ML at index. We do a lot of, I mean, this is all the ML problems that we solve at index. I'm not going to go through all of them

A lot of these are in their second or third iteration. Some of them we've just started, especially the deep learning based problems. And some of them, they're an early version. So what I'll do is I'll pick up two of these and go into a little bit detail so that you guys get context on why some of the pain points we faced and how did we solve them. So this is a problem of classification. So assume that we have the same product coming from four different retailers. In this case, you can see it's from Amazon, Walmart, Walgreens, and Target. So the problem here is we have to categorize them to a single taxonomy in our single taxonomy node

So in this case, it happens to be the LipCare. If you see, the information that we get is the title, the breadcrum, and even the description. So we have to now put this into, I mean, all four of them have to be categorized into this. A single category. So this is a multi-label, multi-level hierarchical classification problem. We have an ensemble of SVMs that actually do this. And we also have, you know, specific models for different regions. The India-based classifier actually uses, you know, word-based embeddings using fast text

The second problem is AdWord extraction. So in this case, you know, you're given this same information. You might be given the description and the spec text also. So you have to extract stuff like, you know, what is a brand? What is the flavor? What is the unit weight? And the pack size. So this is, you know, standard NLP problem. So we have, you know, used CRF-based techniques to solve this for now. But we're also investigating RNN-based approaches. Okay

So that gives you a, you know, flavor of some of the problems we are facing. Let's actually talk about the typical machine learning workflow. You know, I won't go into a lot of depth into this because, you know, every machine learning talk actually goes through this. So I'll try to do this quickly. So, yeah, you start with defining, you start by defining business objective. The first thing you do is, you know, you pull and acquire data. You try to gather data. You look at the data

You explore the data. And then you figure out, you know, what are the features that will help you. Then you start developing a model. And then you evaluate the model. And this process might be iterative. And you see if this model, you know, meets your business needs. If it does, you go ahead and build the production system and deploy it. If it doesn't, you know, this whole process, you know, you go through it again

And once you deploy the model, I mean, you still have to measure, you need to measure to make sure that it's performing the way that you expected it to. You know, is it still meeting the business needs? So you typically, you know, have a human in the loop who's looking at the metrics to make sure that things are fine. And if not, you know, the whole process repeats. Right? I think this everybody is familiar. And everybody, anybody who's doing machine learning, you know, would typically follow this approach. Now, there's another way of looking at this, you know. So this is, there's a TechCrunch article I read, you know, a few days back that talks about this as what is called a machine learning sandwich. So you can think of, you know, three aspects here

The first one is the bun on the top, which is, you know, acquiring data and, you know, transformations. The second one is, you know, the meat, which is developing model and model evaluation. And the final one is productionization. And it turns out, at least, you know, based on our experience, the meat is not in the middle. Right? And what we found out, you know, so we've been doing this for about five years. And what we found out that, you know, typically 80% of the time is spent in the top and the bottom bun. And 20% is actually spent in the models. I mean, is this, is this what you guys also feel? I mean, is this, and I think, you know, even experts, experts also agree, you know, this is a paper written by Google in 2015

So that's, so only a small fraction of real world ML systems is composed of ML code as shown by the small box in the middle. The required surrounding infrastructure is fast and complex. So you see that box there, and that's what is ML code and the rest of the stuff is all the infrastructure surrounding it. And, you know, we, we, we face similar issues. I think there's another reason why this is, this happens. Because, you know, you need different skill sets for each of these three areas. You know, you need, you know, people who can build data, data pipelines, you know, usually called data engineers to solve the first problem. The second problem involves, you know, your, your typical, I mean, your data scientists who know a probability and, and, and, and statistics

And finally, you need app developers, you know, who can take these models and take them to production. And at index, you know, we've, we've gone through multiple iterations of it. You know, initially we tried to look for a unicorn who could do all three and we miserably failed. We tried to pair folks. And most recently, you know, we hired, you know, the team eventually grew to about 15 people who would, data scientists. And they got frustrated because, you know, they were spending a lot of time in number one and number three. And they finally left. I mean, within about six months, you know, the team became, was, was from 15 to zero

So, you know, we, we've actually felt a lot of pain because of, you know, because of these things. And one of the things, you know, we, we, we decided is, you know, let's try to see, you know, from some of the engineering principles that, that techniques we've learned, you know, can we solve one and three. And, so there's a separate talk by Manoj who's here, you know, he's also from index. So he's going to talk tomorrow at 9 50 on, you know, how did we solve this, the first problem. I think, you know, we made some, this, this system is live, it's been there for about two years. And I think we made some serious progress there. So, you know, please attend the talk to understand how we've solved this problem. So my talk is, is, is going to focus on the last bit, which is, you know, taking models to production

So, so to be a little specific, you know, let's actually talk about the, the, the pain points that we actually faced while working on these ML systems. This is a true story. A key employee in the team had to abruptly go and leave. And what happened is, you know, we were not able to reproduce the performance of an existing model in production that he built. And, and what are some of the reasons? The training data that was used to build this model, you know, we, you know, we were not able to find it or it was missing. And the scripts that were, that were there for the pre-processing, I mean, they were not changed. They were not, they were not, not, not in source control. And hyperparameters not known

Hyperparameters are, and these are things that are not learned by the model. For example, if you're solving, if you're using, you know, K and base clustering, you know, the value of K, or if you're using a random forest, you know, the max depth of the tree, the max of the tree or the number of trees, those are, those are hyperparameters. So we didn't know what were the hyperparameters that were, that were used to get the final model working, you know, which, which performed, you know, the best. Has anybody faced this similar problem? Okay. And, and I mean, it doesn't look unique, but yeah. And it takes three months to practice a model. I think this definitely, right, you know, show of hands, how many of you face this problem? Yeah. I think people are just lazy not to, not raising their hands, but

Okay. And the reason that this happens is, you know, because there's a lot of glue code that you have to write. There's custom code developed every time. So we were not able to reuse, you know, some of the stuff that were done by teams working on some of the other projects. And this also means that, you know, frequent updates to models take, take a long time, you know, again, because of, you know, the similar reasons. And another aspect is, you know, heterogeneous systems. So we are primarily a Scala shop, you know, most of our data processing is all in Scala. So we use Scalding and Spark heavily

And we also have built a lot of domain logic over the course of the last five years. And it was very difficult for us, for us to share, you know, stuff between Python and JVM. And we, and some cases we would also use the JVM because for performance reasons. And those are the pain points. But there's also reality. Confidence in test set is not equal to confidence in production. So, you know, you, you would, you know, use different, you know, sampling techniques to actually get training data. But then, you know, you always know that, you know, there'll be some data that your model would not have seen and that you'll encounter in production

So you need to be on top of it to make sure that, you know, you're rebuilding this model. So, you know, there have been issues, you know, where we actually pushed a model into production, but realized a few weeks later that, you know, it's not performing as well as the previous model. Even though, in test set, it was performing well. So, okay. So, I, you know, I actually had a deja vu, you know, when I, you know, started looking into this team. And, you know, some of these problems, you know, seemed very similar to the last mile problem of software delivery. Right. And, you know, so there's this book on content delivery written by, you know, Jez Humble and David Farley

And, you know, I had the fortunate, I mean, I was fortunate enough to work with Jez while he was writing this book. He was a product manager and I was a tech lead on the same product, GoCD, that I spoke about. So, you know, so I thought, you know, can we apply some of the principles from content delivery into machine learning, you know, taking, I mean, taking these models to production. And that's, that's what this talk, I mean, the rest of the talk is all about, you know, what are the principles that, you know, we, we corresponding principles in software delivery that we apply to machine learning. So let's actually go through that. Yeah, before that, you know, I'll just, you know, talk about what is content delivery. So yeah, it's a, it's a software engineering approach that aims at building, testing and releasing software faster and more frequently. The idea is, you know, the time that it takes from a commit, from, from, from, from, the time from a commit making it to your source code till it, you know, till that code goes to production

I mean, you would like to reduce that. You would, you would like to automate this entire process. And you want to do it in a straightforward and repeatable process, you know, that's important from, from a content delivery perspective. So let's look at, you know, what principles we applied. The first one. So content delivery talks about, you know, automation via CI plus CD pipelines. So, you know, a similar principle out here is automation of ML training, evaluation and offline prediction pipelines. If you look at the first pain point, it was, it was about reproducing a training, the entire training pipeline and the model that got created, right? So this, this is that

So how did we go about doing this? So we thought of, you know, trading pipelines more as, as build pipelines. If you've, you know, if you've heard of build pipelines, you know, build pipelines are how you, you know, model the different stages of a software lifecycle. You know, you, you start with a build lifecycle, sorry. You start with, let's say build and you say test, then functional tests and so on and so forth. And then you deploy. So we, you know, we use similar concepts from there. And what we also did was we customized Go CD, which is actually a CI CD tool to, to, to, to make it work for this. And, and how did we do, you know, we actually added plugins to help us with our ML workflows, you know, outside the Go CD team at ThoughtWorks

We are one of the biggest contributors of plugins. I'll actually show you how we did it. But before that, let's, let's, you know, look at, you know, how the pipeline looks. So we have, I mean, a typical training, training pipeline might, you know, have these three, and there would be more, but, you know, this, this, this in, this in essence captures, you know, what usually happens. You first, you know, pre-process the data in, let's say, in a Spark job. Then you would build, I mean, then you'll build a model in, in using Python script, and then you would evaluate the model. Or you would have a slight variation where, you know, you have the model being built in, in Spark itself. Or, you know, if some data scientists want to work on, let's say, using notebooks like Zeppelin or Jupyter, you know, they would like to automate the notebook part as part of their workflow

So I'll just show you a quick demo of, you know, how we, how we actually do this. Yeah, so, so this is GoCD. So what you're, what you're seeing here is, I'll probably just, yeah, so, so this is one of your training pipelines. So this is, this is the first stage called pre-process. So behind the scenes, you know, it's running, actually, a Spark job. Build model, you know, it will be running a Python script. Evaluate model, again, a Python script. So there's an example that, that I just created

But this is a real life, I mean, this is a real ML training pipeline that we have at Index right now. So you can see there's a label here. Yeah. We'll, we'll come back to this again. The next principle is, you know, having a source code and artifact repository for reproducibility. I mean, this is, again, that, what, what we talk about in software. A similar analogy here would be data and model repository for reproducibility. The data part, again, you know, Manoj will talk about in his talk

So I'm going to focus on the model repository. So you can think of a model repository as similar to an artifact repository like Maven or Ivy. We actually heavily borrowed, you know, stuff like the directory structure, the versioning, the semantic versioning, as well as the fact that you have snapshots and releases from there. Even the publishing of models is, is, is something that, you know, we, we borrowed. And we implemented clients that would allow published, that, that published model for most used frameworks. We built, we built clients for scikit-learn and spark MLlib, MLlib and Keras. So I'll actually talk about, you know, spark MLlib a little bit. So, so every ML algorithm in MLlib has, implements this interface called MLwriter, which has, which has, gives the ability to save and load your model, model and model metadata

So we, we, we, we, we use that, you know, for this purpose. It's opaque. I mean, it, it, it stores the model metadata in a JSON format and, and it stores the data itself in Parquet. But for us, for our use case, I mean, it worked really well because, you know, we would, we would, you know, we were just using it, you know, we were using Spark for the entire end to end, at least in, in, in the cases where we used MLlib. Yeah. So, so for a model, you know, what, what does it store? A data is, you know, we, we, we would store that data in S3. In different formats. For example, for MLlib, Spark A, for scikit-learn, pickle, and for Keras, we use the H5 format format

And we also would store the metadata for a model. So what, what is, what is, what is, what is, constituted in the metadata? What were the training validation and test data sets used? In this case, it, it keeps a pointer to the, to the data platform that, that, that, that, you know, we built. And then what were the hyperparameters used? And then finally the evaluation metrics. So this goes in, this is what goes into the model repository. And then, as far as automation goes, once the model, model gets evaluated, we use, we use the model repository client to publish the model. Apart from that, you know, we also need to think about model promotion. I know this is similar to build promotion. The reason being, I mean, you, you want to tag the latest good version that needs to be deployed because, you know, not, you know, every, every model version, not every model can be, need or can be promoted because you might be working on some experimental models that you don't want the downstream systems to use

And you also want, you don't want models that fail a specific, you know, test of performance set to, to go through. So what you do and the other advantages that you can, because it's just a tag, you can roll back to a previous version. So, you know, taking the same automation example, workflow example. So there is this new stage that we had called promote model. And this is a manual stage. Let me just give a demo of that. So if you look at this pipeline here, so there's the published model here, and this stage is actually a manual stage. So you can go ahead and trigger this

So in this case, what you're doing is you're, you know, it's calling an API to tag, tag this particular version. In this case, let's say version three as the, as the latest good. And let's say if you want to revert, you can just go back and, you know, okay, let's, if this finishes, you can, you know, this will be available and you can redo it. So this is how you can, you know, roll back a specific version. Okay. Yeah, moving on. The next principle is, you know, using containers for microservices. I mean, that's something that's, that's pretty common now in the area of software

So similar thing here is, you know, model containers for model prediction microservices. So I'll, I'll talk about that. So, you know, we started thinking about, you know, model containers. So model container is, you know, a container is your unit of deployment, you know, so, so we would, you know, we would have a single, a single model hosted in a container and that will be used for predictions. And this, you know, container would, would expose an API for prediction and it, and, and, and it is dockerized. And these containers can be replicated to handle, handle scale, right? I mean, that's, that's the true promise of containers. You know, you don't have to worry about scale. And the way that we do it is, you know, we do it using two microservices

So, so we have a Scala front end, which handles, you know, pre-processing. The reason we do this is because a lot of our, you know, domain code is, is in Scala and we don't want to, you know, rewrite it in Python. So that's why, you know, we do this. And, and we use Python because the ML ecosystem is much better in Python and most of the guys we hire, I mean, they're comfortable working with, with Python. I mean, you might also ask, you know, why did we not just, you know, export the model from Python and use it in Scala? Multiple reasons, you know, we tried looking at multiple format, formats that are available. So PMML, the issue there is, you know, it doesn't work well, very well with ensembles. But I've heard, I mean, just, just yesterday I heard that Stripe is working on some interesting stuff where they've written custom encoders that can, you know, read Python, Python serialized objects. So that, that definitely seems, seems interesting

And there's also, Apple's core ML tools. And those are a couple of things that, you know, are next steps for us to look into. But at that time, you know, we, we, we found it better to, or simple to just use two services, you know, Scala for the pre-processing and Python for the model part. Loading the model and, yeah, so, so Python loads the model and exposes the predict on the model. You can also predict in batches for better throughput. And, you know, so all the ensemble of models is all abstracted within the Python, within the Python service itself. And the Scala microservices, you know, delegates the predict and predict batch functions to the Python microservice. And this is how it looks like

And if you talk of automation, I mean, some of the steps you would do is, you know, you would create a Docker image for this entire thing. And you would push this to Docker registry. The next step is model deployment. So we actually had two modes. We had a, you know, I spoke about the fact that, you know, we have used Lambda architecture at index, right? So we have a batch system. We have a batch pipeline and a real-time pipeline. So we have a batch pipeline, which means that most of these models actually need both offline and online mode. So for the offline mode, what we would do is we would package this model containers into an AMI

So we are completely hosted on AWS. So AMI is their, you know, is the Amazon machine image format. And what we would do is, you know, this AMI would start as part of the containers, as containers when you start your Spark and Hadoop clusters. So we would, whenever we bring up a new, during a batch pipeline runs, we will bring up Hadoop and Spark clusters. So these containers will already come up with the AMI. And they'll be running on each of the executors and task stackers respectively. And within the Hadoop or the Spark service, we would actually call the local Scala service for the prediction. The online mode, you know, we, we, we've been using Mesos and Marathon for about three years now

So we have a bunch of services on Mesos and Marathon. And more recently, the last couple of years, actually probably a year, year and a half, we've been using Kubernetes. So, you know, we, we've actually pushed these, you know, the model containers both to Mesos and Marathon on Kubernetes, depending upon the, on the machine, on the, on the project that we're working on. And the scaling is actually taken care of by, by these clusters. Yeah, so this is how we do model deployment. Okay, the final principle is, you know, how do you do A-B testing? So, so in, in, when you, when you talk about for software, you typically use concepts like you know, you know, you have a set of instances, micro services running, if you want to push a new version. So you, you will actually push, let's say 20% of the traffic to these, these new services. And if things look good, you know, you, you slowly, you know, push 100% of the traffic

So we use something called request shadowing, similar stuff. I'll tell you why we, why we could not use canary releases. Yeah, so the reason is because, you know, we can't use multi-armed banded testing, you know, which is the typical, you know, principle, I mean, it's typically what is used for testing, model testing, A-B testing of models. The reason is that, unlike, you know, in situations like, for example, CTR, where you can actually push, you know, 20% of your traffic through, you know, to the old, to the new service. And you can, you can see, you know, what is the click through date. You can actually figure out what the, what is the payout. In our case, it's very difficult to figure out the payout because we don't have, you know, we don't have an end user who's actually using, using, using it. And then, you know, we don't have clicks that we can measure, for example

So instead, we use the request shadowing pattern, where, you know, the idea is, let's say, you have an existing, model that is running, and there are 10 instances of the model running. So, and if you need to push a new model, we'll, we'll, we'll bring up another instance. And we'll make sure that, let's say, you know, the nine instances will still be running with the, old service, but, but the 10th instance, and this new, and the new instance we need to pick, what we'll do is, we'll do request shadowing. So every request that will come in, it'll, it'll go to, it'll go to the old service, as well as, it'll go to the new service. But the output will only be taken of the old service. And what we do finally is, you know, we, we, we try to look at both the outputs. You know, if, if they're same, then we don't need to do anything. But if there is, there is a difference between the two, then we do a delta of that and send it for spot checking

And for offline, we only do, you know, deltas in spot checking because, you know, we, we control when the pipeline will run. So if the old pipeline ran with the, the old batch pipeline ran with the old service, you know, we have the data, we'll run the new pipeline and then do a join and compare the data. But spot checking is still, you know, something we do. And for spot checking, you know, we actually have built an in-house data talking tool. We, we of course tried, you know, using mechanical turg, but, you know, there's a lot of domain expertise required in our case. So we ended up building our own tool, which, and there are workers, you know, who actually use this, you know, use this tool and then do the spot checking for us. So this is, you know, an example. So one task is, you know, you need to figure out, this is from the matching, matching service

So you need to figure out if these are the two same products or not. So you have images and the title and category information. So, you know, they just have to say yes or no. And this is what we'll use for spot checking. And this is another example where it's a classifier in action, where, you know, all these have been predicted as LED and LCD TVs. So during the spot check, we need to say yes or no. So yeah, so this is actually my talk and what's planned in the future. So I think there's still a lot more stuff to be done

Support for deep learning is not yet there. I mean, we have a couple of models in production, but, you know, we haven't followed a lot of these approaches. I mean, we need to, you know, think through that as a first class thing. We want to build model repository visualization so that, you know, you can see how the models metrics are evolving over a period of time. We also want to add more plugins in Go CD to better support ML workflows natively. You know, it's right now it works, but it's not. I mean, there's a little bit of, you know, boilerplate that you have to write. Apart from that, you know, if folks are interested, you know, we can definitely look at open sourcing the models serving repository and clients

This is again, you know, feel free to talk to me to see if, you know, if you folks find this interesting and we can, you know, look at that. Yeah, so we are hiring. So it's just a couple of days back. We opened a Hyderabad office back in India. So if you know anybody who wants to move to India and is looking for some interesting work, you know, let me know. And, you know, we are a big supporter of open source. You know, we use a lot of open source tools and love contributing back to the community. So Manoj here is one of the biggest contributors of open source at index

Feel free to talk to him. And please go and check out OSS.index.com. Yeah, that's it. Thanks a lot. Thanks, Josh. Anyone have any questions? How do you think your index will change when everybody starts using RFID tags or products? Sorry, I... Are you using any RFID tags to identify products and how will that impact your products? We have a whole bunch of data coming along with us. I mean, that's a use case that's not come up, but I think there are enough challenges we still have to solve in existing..

Yeah, I know, I know. I mean, that might change a lot of things, but I mean, it's not a use case that we've seen yet. Yeah, so I'm here. And please, if you want more details, please feel free to talk to me. Thanks a lot.