Scale By The Bay 2018: Milind Bhandarkar, Hadoop Future in AI World
Recording: Scale By The Bay 2018: Milind Bhandarkar, Hadoop Future in AI World
There's more functional talks going on in there now. So not a lot of people are interested in the Hadoop or even AI looks like, you know. So that's fine. Just to sort of have a quick survey before I get started. So when did you first hear about Hadoop? Maybe not, you know, worked on Hadoop or worked on top of Hadoop, but hear about Hadoop. Okay. Before 2010? One, two, three, four. Wow
Okay. Before 2006? That's a trick question because 2006 is when the project was called Hadoop. So it has to be 2006, not before 2006. Anyway. So what I will be talking about is, so I'm Milind Bhandarkar. I've been associated with Hadoop even before it was called Hadoop or even big data as the word basically hadn't become common then. Right in about 2005 when I started working at Yahoo in a project that essentially became Hadoop. So I'll be talking a little bit of history about that
And since 2008, I've been giving, you know, sort of evangelizing Hadoop. So it's been about 10 years now that I think the first Hadoop tutorial outside of Yahoo was conducted at ApacheCon in New Orleans. So it's been about 10 years. And one of the things that's been common for the last 10 years is that the future of Hadoop always kept on changing. Right. So I've been giving these talks now, I think about from about 2008. And each of these talks, the variations of what people said Hadoop will evolve into, or what Hadoop is good for, that changed considerably. So I'll be talking a little bit of history about that
Right. All right. So we all know the birth story of Hadoop. There was 2003, 2004. There were two papers published by Google about how they built their search backend infrastructure on Google file system, as well as the Google map produce distributed processing framework. Right. Around the same time, 2003, Yahoo had acquired a company that was building the search backend for the web called Ink2Me, which had started in 1996. And another company called Overture, which was placing ads on top of in those search results
Right. And combined both of those in around 2004, Yahoo launched something called Yahoo search. Right. Ink2Me was the backend for Yahoo search and Overture was the ad placement platform. Right. So at the core of the Ink2Me search backend was a was a process called W or web map. Right. So essentially, think of it as a page rank, you would crawl all the web pages that you know about all the public web pages, you basically form a large graph out of those web pages, and you analyze that in order to determine the relevance
Right. So this web map backend was a project that was launched in 2005. Even with less than a billion public web pages used to take something like six weeks to build. Right. And the web was growing rapidly at that time. In like two, three years, it had grown almost 10x. And so the goal at Yahoo search backend was to actually process this entire web map and index these pages and compute their relevance in one week. So the project was called W1W
So that was launched in 2005, somewhere in the summer. And that's the project that I had joined. So the processing framework or the infrastructure framework for building that web map web map at that time was called Dreadnought. It was built sometime around 2000. And it was, you know, from 20 machines, it had scaled to about 600 machines, but it had reached its end of, you know, scalability. So as part of this W1W project, we basically started writing, you know, five, six people, team, we basically started writing this new infrastructure framework called Juggernaut. Okay. And as part of that Juggernaut, it was basically modeled after the Google papers
So Juggernaut file system and Juggernaut map reduce, those were the two components. And as a back scheduling engine, we had adopted a project from University of Wisconsin called called Condor. Right. So at the end of December 2005, we actually had a pretty nice prototype working on about 100 machines that we had borrowed from Yahoo search. And at the around the same time, we discovered that there is an open source project in the Apache Lucene ecosystem or Apache Lucene project called a backend system called Nudge. And this person called Doug Cutting was developing along with one grad student called Mike Caffarella. He was developing this thing called Nudge distributed file system and Nudge map reduce, NMR and NDFS. Right
So we contacted him and we convinced him to actually separate it out from the Nudge project, because Yahoo had a proprietary search engine and by, you know, our lawyers would not let us contribute to an open source search engine. However, the Hadoop part of things are sort of the Nudge distributed file system part of things. Those were not related to any search technology. And therefore, if they were separated out into a separate project, then Yahoo would be or Yahoo engineers would be able to contribute it. And that was essentially how Hadoop was born. Right. Hadoop that the the map reduce and the distributed file system was separated out from Nudge. Okay
So in the three, four months that we started working on it from 20 machines again in the, in the beginning in January 2006, we made it to scale to about 600 machines in the, in the next three, four months. And around April or March, we launched the first 600 machines cluster called Kryptonite at, at Yahoo. Why 600 machines? Because these were the 600 machines that were three years old that the Dreadnought project was using and they were end of life. So they just let us borrow that in the same data center that they were sitting in and we started using it for Hadoop. Right. So what was the Hadoop future looking like in 2006? So at that time, it was a firm belief at Yahoo that Hadoop as a backend technology that allows us to massively process the web data will allow Yahoo to win the search engine wars that were going at that time essentially between MSN, Yahoo and, and Google. Right. Um, we know how, how that all worked out
Uh, but differently, you know, Hadoop at that time, what, what was Hadoop at that time? It was essentially HDFS, Hadoop distributed file system plus this MapReduce programming framework, right? That's it. All right. Um, well, Yahoo might not have won the search wars, but, uh, I think we, we ourselves in the Hadoop team learned a lot of lessons by deploying Hadoop at this massive scale. By the time I left Yahoo in 2010, December, we were running Hadoop across three data centers in 45,000 machines. Okay. Having 70 some odd petabytes of data. All right. So the lessons that we learned is that multi-tenancy is something that you cannot graft after the design of the fact, after the design of your, your system, you have to design it from the beginning to be a multi-tenant system
Right. And so when we were running these clusters, we were basically at least 1500 Yahoo, uh, engineers and developers used to run or use these systems continuously. Right. We had to design it from the ground up, uh, to be a multi-tenant system. The second thing is that too much focus on performance kills the agility, right? You should design your systems to be agile rather than extracting the bare metal performance out of this. Right. The third one, uh, and this actually is now with the cloud and everything else is actually now we have realized it a lot more is that the, model of infrastructure, which is procuring the hardware and then using it for something, uh, right. Uh, provisioning the hardware is much simpler and much easier than going through this entire procurement cycle
Right. So just to sort of a sidebar here, uh, on every Friday, uh, morning, actually whole day, uh, in Yahoo, there used to be, uh, a meeting, uh, called the, uh, of a committee called, uh, hardware requisition committee or HRC. Right. Any, any old timers Yahoo's here? No, probably. Okay. So, so this is then an interesting case because you would know about this. Um, this HRC meeting was actually led by, uh, the founder of Yahoo, co-founder of Yahoo, uh, David Filo, right. And he is, his staff used to sit around the, uh, the table and anybody who needed any new hardware to be deployed or to be procured used to appear before this present their use cases, present what kind of hardware they needed and had to justify, uh, why they needed that hardware
So by law in Yahoo, uh, every machine that was running in Yahoo, uh, by the bootloader itself, it had actually created a user account for David Filo. So you will basically go to David Filo and say, Hey, I need these 10 machines because I want to run this new ad pipeline or whatever it was. And he basically said, okay, where are you running this now? And then you would basically give him the host names, uh, for, for these machines. And he would actually log in into that, do a visor and, uh, or just the Yahoo version of SAR and basically say, okay, but these machines are being utilized currently about 10 to 15%. Why do you need 10 additional machines? Right. I have seen managers, like grown managers cry in that meeting, trying to justify why I need 10 more machines for running my workloads. Right. When Hadoop came along, right
I think the, the one thing that we sort of decoupled, uh, uh, application developers or the data pipeline developers from is this HRC. We went in there and we basically said that, Hey, the look at the Hadoop adoption curve, it is growing at this rate. We need 8,000 machines within half an hour. We'll be out of there having, you know, got a signature of, uh, uh, David Filo. And then these users then will come to us, say, can I have a hundred machines out of that? And that would be much easier. Right. So provisioning those machines for other users was much simpler than everyone going and procuring those machines. Right
The second thing is rather than the prevalent practice in Yahoo, where the, the platform team actually developed the software and handed it off to the user teams and the user operations team deployed it. We ourselves who were the developers of Hadoop within Yahoo, we ourselves deployed that and, and, uh, provided that, that as a platform to the rest of Yahoo, right, which basically meant that we could observe what people were doing with our software. And when, even if we designed it to be like this large scale data processing layers across multiple petabytes of multiple hundreds of terabytes of data stored in HDFS, people would always run some sort of a weird, you know, job that would either take down the cluster or maybe, you know, miss make others miss their salaries. But since we were actually running these clusters and observing what people were doing, we could then go back to them and basically say, what is it that you need to run that is not normally supported by the software? And we added those features into the software. Right. So, so these weird use cases we used at a lot of learning, experience. The, the fourth thing, because we built this thing open source, the fifth thing that we did is basically we had a lot of academic collaboration or collaboration with a lot of universities, CMU, uh, you know, you know, Berkeley, MIT, et cetera, et cetera, and onboarded their grad students to use one of our publicly available clusters, which was literally sitting in one of the buildings, uh, parking lot in, in, in Yahoo in a, in a large trailer. So that cluster called M45, we opened it up to the students and we could then even observe what the students were trying to do
Grad students were, how they were trying to use this, these machines or these, uh, infrastructure, and then, uh, you know, learned a lot from that. Okay. So years passed. I left Yahoo in 2010, joined a company called, uh, uh, Greenplum, which was, uh, launching into their Hadoop strategy. And as these big, and this was the division of EMC at that, uh, that time. So as these big enterprise vendors got attention, uh, or, or started focusing on Hadoop, that resulted in a peak hype of Hadoop and the distro wars that resulted from that in about, uh, 2011 and 24 or to about 2014. Okay. So how many people have seen these headlines saying that, Hey, half of the world data is going to be landing in Hadoop or getting processed in Hadoop, every application or whatever, 65% of the analytics applications will come embedded with Hadoop
So all these, you know, analysts and the enterprises basically started focusing on Hadoop. And as a result, what Hadoop was good for or what it was being used for that started getting stretched left and right to address all sorts of different use cases. Right. Um, all right. So, so then I, if you have, if you have seen some of these slides on the web, uh, you know, uh, these, these are part of my earlier presentations about during that Hadoop hype period, I've just reused some of those pictures there. So we basically went out there to the enterprise and basically said, Hey, your analytics process itself is, Hadoop is going to disrupt that analytics process. Your current analytics process basically takes a lot of data, goes through a costly batch ETL cycle, drops most of that data, pushes some of that structured data into your data warehouse and, uh, this thing. And as a result, you are losing 80% of your data and losing, you know, uh, uh, valuable business insights, right? What Hadoop can make you do is right now is basically, uh, you can, you know, first put all this data in the Hadoop infrastructure
And instead of the ETL, you basically do extract load and transform ELT, and then you can start using some of those in your, uh, the, the structured data sets in your EDW, as well as your data marks, right? However, the second, the third thing that everybody then started doing is that, Hey, why take this, use Hadoop only as an ETL offload or an ELT offload? Why can't you just perform analytics directly on top of that, right? So between 2011 and 2014, the biggest focus was on making use of Hadoop as an analytics platform for the entire company, a multi-tenant analytics platform, right? And for that, we basically needed some, uh, uh, the lingua franca of, uh, of, uh, analytics, which was SQL running on top of Hadoop. MapReduce itself did not, uh, you know, make sense when you wanted to have business analysts give access to your Hadoop cluster. So that is where a lot of SQL development happened, right? Whatever might have happened to that promise, I think one thing that became clear is that the data economics was, or data analytics economics was forever altered by Hadoop, right? When we basically started implementing analytics workload on top of Hadoop, the, the, the traditional data warehouses, the MPP, EDW, et cetera, exadata, teradata, all these things used to cost several tens of thousands of dollars per terabyte of processing of data, right? That came down, uh, uh, significantly because the pressure that, Hey, I could, uh, you know, offload some of that analytics to Hadoop actually started growing and these MPP database warehouse vendors started coming down in cost, right? And as a result of that, there was basically this, this new, uh, when Hadoop people basically suddenly discovered that, Hey, there is something exists that basically every analyst uses, which is SQL, right? And Michael Stonebreaker, the Turing award winner, uh, and, and, uh, a frequent critic of Hadoop essentially called this no SQL movement as the not yet SQL movement. You start with basically saying that I am not going to, you know, be providing the, the SQL interfaces. I'm just going to give you the Java APIs. And as the technology matures, you basically suddenly realize that, Hey, I want my business analyst to use my data stores, my data processing platform. Therefore I'm going to be adding Hadoop, right? Uh, the SQL, right? So in those, uh, I would say about, uh, Hive started a bit earlier, but in those three to four years, there was something like, uh, uh, 10 different SQL engines, distributed SQL engines directly on top of Hadoop. So the, the SQL on Hadoop essentially started, uh, proliferating during that time
Okay. So what was the Hadoop future then in 2014? Everybody, uh, and their uncle basically said that the traditional enterprise data warehouse that we, uh, as we know it, Hadoop is going to kill that, right? And in this case, what was Hadoop? Hadoop was essentially the Hadoop distribution, uh, distributed file system, a scheduling, uh, framework on top of that called yarn that had come up in Hadoop 2.0. And on top of that, the SQL of the various flavors of SQL on Hadoop, right? That was essentially what was, what was Hadoop. So things were going really well, uh, but 2014 basically big, uh, you know, started, you know, the, the became a turning point for the Hadoop future as we know it, right? So the future was disrupted in 2014. Any, you know, guesses about what happened, what significant things happened in 2014 that might have disrupted the future? Spark essentially. Well, people talk about spark a lot, but I think spark still became part of the, yeah, Hadoop distributions adopted spark and essentially integrated it within, Hadoop itself, right? By changing the definition of Hadoop somewhat, right? In my opinion, one of the biggest thing that happened in, in 2014 is the public and private clouds really started accelerating, right? I mean, Amazon AWS actually started back in, uh, around the same time as Hadoop started around 2006 or something like that with just S3 and EMR. But then over the years, they started adding these services, uh, left and right. And essentially, uh, I remember in 2014 reInvent, I think, uh, it was the CIO of, uh, Goldman Sachs or somebody like that who appeared at the reInvent and on the, on the keynote stage and basically says, Amazon can do infrastructure much better than any of us would ever be able to do, right? And that is when actually, and, and the, they can make this infrastructure, uh, much more secure than any one of us would be able to do, right? And that's where the public's perception of, of these clouds basically started changing
Other otherwise, before that people were, you know, always giving reasons of security, lack of control, lack of SLAs and lack of services and very immature and all those things. 2014 changed all of that, right? Essentially, the infrastructure as a service became the new hardware, right? You started adopting, uh, uh, these large scale virtual machine farms as your own hardware, right? Obviously the public clouds were there, but even in the private clouds, vSphere, VMware and, and, and OpenStack started, you know, making the infrastructure as a service reality, even in the private clouds, right? Obviously the, the, the advantages we all know, very easy provisioning. You basically make a rest API call to basically a provision of virtual machine. It's much more scalable than you would ever be able to build it yourself, right? It is elastic in the sense that you could throw, throw away the resources when you don't need it. And it's ubiquitous. You could basically see, you know, across all, all, all regions of the world, you have the, the infrastructure as a service available, right? And increasingly it came bundled with the data storage and analytics services, analytics as services. I would basically say the cloud data fabric basically work was, beginning to get so scalable. I mean, exabyte range was no longer, you know, the, the, the, the, top goal for any company
Exabyte scale storage was available. Just swipe of a credit card from your object stores, et cetera, right? So all the services that were built around this exabyte scale storage, I mean, multi exabyte scale storage to integrate ingest from various data sources, you know, ability to rapidly analyze these data sets by the scale out infrastructure as a service platform. All those things started coming around, right? So those were the two major things, or I would say cloud was the first major thing that essentially disrupted the future of Hadoop or the trajectory where Hadoop was going, right? And now we increasingly start seeing after the, for the last couple of years, the modern AI based workloads or ML based workloads are the, are the second disrupting factor that basically are, you know, contributing to this trajectory change of Hadoop, right? And I'll, I'll focus a little bit on that. So starting with a joke, you know, I don't, I don't, you know, endorse this statement. I think a lot of AI workloads are real AI workloads. These are not just select followed by group by clause. But the, but the, I think, I think the notion behind that is that, hey, data is still extremely important for these new AI based workloads, right? So that's the notion that basically comes, from there, right? So in, I think this was last summer that somebody published from, from Google Research and CMU, somebody essentially revisited a paper that Peter Norvig had written 10 years ago called Unreasonable Effectiveness of Data. So 10 years ago, Peter Norvig's contention was that, hey, we did all these, you know, natural language recognition kind of projects from the first principles saying, how's, how's the natural language constructed? But then literally what changed the, the, the, the speed of, you know, progress in the natural language processing when a lot of this natural language data became available and we could process it at a really, really fast rate, right? So that was 10 years ago
Now with the new techniques, the deep learning techniques, they revisited that, that paper and basically say, is the, is the data or the scale of data still effective in, in these new deep learning kind of workloads, right? And what they found is that, you know, this is, I think on a number of videos that they had trained on the, the effectiveness of these deep learning models obviously increases when more and more samples are fed into that, right? I mean, that, that sounds sort of almost true. But if you now look at the, the, the scales on both of these graphs, this is the X axis is actually, X axis is actually log scale, right? So in order to get your deep learning models more and more effective, you need to throw more and more data, like an order of magnitude more data added, right? So it means that, you know, the, the data, the, the scale of data is still very much needed in order to, uh, uh, uh, be, you know, true in this, uh, uh, in this, in this AI and deep learning kind of era, right? So why is it that people are basically saying, uh, especially when you go to Sandhill Road, uh, that's where the lot of VC community is, uh, why are they saying that the big data is passe? We don't need big data anymore. AI is the new hotness because these two things are very closely related, right? So big data is still important, right? In the AI world. So why aren't people considering Hadoop as the platform for building your AI applications or AI, AI models, et cetera, right? So now when I stand here in 2018, so what does, what does it mean to be Hadoop, right? It's, it cannot just be a SQL on a Hadoop processing engine. It cannot just be HDFS and, and yarn because now we have a cloud infrastructure and, uh, instantly scalable with a large cloud data fabric, et cetera. What Hadoop is becoming today. And if you see a lot of development in the Hadoop 3.0 world, a lot of these are essentially becoming standardized API for building and managing your AI and analytics workloads, right? So things like TensorFlow on yarn, TensorFlow on spark, you know, all those things are getting integrated. You are given now a, uh, a storage substrate, which is this large cloud data fabric, and you are given a scheduling and we'll talk a little bit about that, uh, a scheduling framework, which is a reference implementation for these larger, large scale analytics, uh, workloads and AI workloads
Okay. So what has Hadoop become in, uh, in, uh, 2018? Well, we were talking about Hadoop distributed file system. I think HDFS gave a, a reference API or, or a standard API for all these different cloud storages to be integrated to the, or connected to the analytics workload. So HDFS has evolved into being an API called, uh, uh, it's CFS essentially Hadoop compatible file system, APIs, and those APIs are implemented on top of all the object storage that you might want to think of, like, uh, uh, Azure object storage or, or, uh, uh, uh, Google blob storage or, or S3, right? A lot of these analytics applications are accessing those object storages, blob storages, uh, from, from, from the, uh, Hadoop compatible file system API. The second part, which is that you had a static compute, uh, resources that were scheduled by yarn or something like that. Um, so those are now basically, uh, becoming containerized, right? So we basically see now, even inside of yarn, there is a support for Docker. And, uh, uh, there, there are a lot of projects which are essentially trying to move that, uh, uh, container orchestration into Kubernetes, even for these analytics, uh, applications, right? One major change that we have started seeing is the, the deconstruction of yarn, and I will talk a little bit about that later. But what yarn was, was a demand side scheduler, right? Or what, how this, this was designed in Hadoop was that you would have a fixed set of resources, let's say a thousand machine cluster
There are only a fixed number of compute cores or compute resources out there. And yarn would be a scheduler for demand scheduling, which is that there will be a lot of workloads coming at that particular cluster. And yarn would basically, uh, uh, schedule those workloads in a, for a fixed set of compute resources, right? On the cloud, this model is completely flipped right now, right? On the cloud side, when the workloads are submitted, you don't have to schedule them within the same resources that you own. You can actually grab new resources in several minutes and you can schedule it there, right? So now the scheduling is changing from demand side scheduling to the supply side scheduling, right? Then the, the, the, the, uh, objectives of this scheduling are very different. The demand side scheduling is basically reserved for, or optimizing the utilization of your compute resources, right? Uh, you would basically say that, okay, how do I pack in more and more workloads in the same set of resources to increase the utilization or a percentage utilization there? The supply side scheduling is basically, uh, uh, trying to optimize the SLA's for your applications or for your workloads and the cost for running those workloads, right? So that's the biggest change that we have started seeing because of the cloud came along there, right? The, the fifth thing is obviously these workloads are going to be strung together by some sort of a workflow, uh, uh, workflow, right? And that workflow models, rather than having a single static workflow scheduler or something like that, that is moving to become more and more like a serverless workflow, right? When you have an, and, and earlier talk in Netflix, which was basically, you know, you, you declaratively specify your workflow and then at the end of the completion of every node, it will send an event and that event, there will be some listener on that event and that, uh, uh, what to, what to execute next, right? That can be done in a completely serverless manner, right? So you basically are started seeing serverless workflow. So what is it that still tries together, ties together all these, these, uh, Hadoop components? Well, the higher level abstractions have emerged, which is again, we started with files, but now we have gone on to tables again as a higher level abstraction or table pods, I would like to call them. And all of those are described using a single metadata service, right? So metadata service is what now tries ties together all of these components inside of Hadoop together, right? So again, going over, you know, each one of these things, the computes are getting dominated by containers and orchestration frameworks. Kubernetes is everywhere, right? We waited for at least three, four years for somebody to emerge as a winner for this container orchestration
And you talk to anybody, I think everybody says that Kubernetes has emerged as that winner. Um, uh, as I said, Yarn is getting deconstructed from becoming a demand side scheduling algorithm or demand side scheduling to become a supply side scheduling. And which basically means that all the internal, uh, uh, sort of features of yarn are getting externalized, right? Workforce scheduling, how do you allocate resources? How do you manage those resources? How do you keep isolation between various workloads running inside those resources? All of that is getting deconstructed, right? And all of the compute is now assumed to be not only logically, but physically separated away from the storage, right? Because that's what the cloud storage model is shared storage separated away from the computer resources. The storage itself is moving to be massively scalable individual tiers, right? So you have started seeing the cloud offerings. Also the storage offerings are becoming premier tier, standard tier based on the throughput and the latencies that you are getting from there. And each one of these tiers is massively scalable, right? You're talking about the, the exabyte scale object stores, then, uh, distributed file systems, which are layered on top of those and all the persistent volumes that you need for your current workloads. All of those are again massively scalable and in a, in a tiered manner, uh, based on the cost and the performance that, that, uh, you expect from that. From the lower level abstractions like files and directories, we have now started moving to the higher level abstractions
Most of the data workloads in any case happens on, uh, even, even if they were happening on the, file system, they first brought in the file system, parse that file, uh, uh, brought in the file, parse that file, created structured representations of the data in that file, and then started compute on top of that. So why not just work, uh, directly on top of those higher level abstractions? And as a result, the metadata services are the ones that are stringing together all these higher level abstractions. What we have started seeing now, uh, and they are actually becoming available on the cloud storage first, is very large and dense, uh, uh, both volatile as well as non-volatile memory services, right? And, uh, that basically then tends to get used as a persistent, uh, caches for the computes because, uh, the, the, compute and the storage are being separated in the cloud area. And therefore, we basically see, in order to reduce the variability of the storage access, we see a lot of these caches being used on the compute side, uh, with this large, large volatile indexes, right? So obviously, uh, the, the, the 800 pound, uh, uh, elephant in the room, we used to have, uh, three elephants in the room. Now there are two because Cloudera and Hortonworks essentially, uh, uh, merged. Um, so obviously I think, uh, uh, this is a great news for the, for the Hadoop ecosystem in general, because now they have to worry less about which vendor do I work with, et cetera. And, and two vendors combined will actually, you know, build more end-to-end data workloads in the cloud. That's my personal feeling
Obviously, uh, uh, people in the industry are now talking about declining, declining influence of, of Hadoop. Let's see what happens. But I think, I think the future is pretty exciting for Hadoop as, as, as long as it focuses on these modern AI and modern data analytics workloads. Okay. That's pretty much what I have. Do we have time for questions? Yeah. Go ahead. So did you talk about, you talked about metadata services, but I haven't seen concrete metadata services at the time
What are your... Right. So, so, uh, the system, I, I, I tend to differentiate between the system metadata services and call it the business metadata services, right? The system metadata services is basically for the computations to, uh, know about how the data is organized, what it's contained in that data, uh, et cetera, et cetera. Right. So I believe one of the, uh, again, Hadoop defines the, the, the, standards. I think the high metadata service has emerged as the standard interface layer, at least to your data sets, right? If you basically see, uh, even if your data is stored on, on a, uh, S3 or on, uh, remote, uh, local file system or this thing, it will be accessed through the high metadata store. Even the, uh, cloud providers own metadata services, something like Amazon glue, for example, they have high compatible, uh, metadata, uh, uh, API layer, right? So that has emerged as a standard. So what remains of Hadoop from the last five years is essentially that high metadata server right now
But is it really being worked upon new things coming to those? Uh, yes, yes. So, so, so, so, so a lot of, a lot of work is actually going on in terms of adding more and more capabilities, not, not to so much to the system metadata servers, but if you see other developments like, uh, integration of this high metadata service in Apache Atlas, which is the, which is essentially a metadata service, right? But it's a business metadata service that tracks lineage and that tracks the usage for collaboration among data scientists and among other analysts. So that's the second part, which is the, uh, business metadata services. Uh, in the last three months, uh, we basically have started hearing about an open, uh, business metadata services, uh, uh, project called, uh, Egeria or Egeria. I don't know how that is pronounced, but that's being led by a consortium called the ODPI consortium, the open data platform initiative. So that's on the metadata side and Atlas is basically a reference implementation for that, right? So we have started seeing those, the emergence of the, the open metadata services in this, in this area now. Yeah. Yeah
Go ahead, Kathir. So one of the, uh, Hadoop is the data lake. Uh, hmm. So now the more and more I see the market is, uh, deconstructing data like specific, uh, workloads. Yes. Yeah. Correct. Well, uh, what will the future of Apache Hadoop? I think Apache Hadoop is going to remain an API, right? And a reference implementation
Uh, but really the, the yarn is the one that actually proliferated these, these kinds of, uh, workloads, right? Back in the days, 2005 to I think 2009, we basically used to have something called Hadoop on demand, right? And the, the Hadoop on demand was actually a misnomer because it was MapReduce on demand, right? So we basically, the way we had implemented this architecture in, in Yahoo was a single HDFS substrate, which will contain all the data, but then different groups of people, right? Different teams would essentially create 150 node cluster here, 160 node on the same machines, but those will be sort of the MapReduce clusters. Those were separated MapReduce clusters, right? Then came Yarn and then basically say, okay, no, single unified scheduler across the entire, entire cluster and all the use cases will be implemented on top of that. And Yarn is the one that actually facilitated move away from MapReduce for all these SQL engines deployment of long running services, all those kinds of things, right? So I think this is the logical continuation of that, which is now that the S3 and all these massive object store basically become the data lake, as in data resting place for that. And then you have all these tiny, you know, clusters coming up on demand and accessing that the, the S3 data lake, right? So then you basically load your data into S3, point Snowflake to that, and that becomes your data warehouse. You load data into S3, point, you know, some sort of a key value store on top of that, and that basically becomes your real time operational platform. So the data lake as a resting place and of data and data lake as the processing part, those have separated, but essentially the shared storage part still remains there. Yeah. Any more questions? Cool
Just from that, I think you are saying most of the research now is happening on cloud, not on optimizing the clear material. That is absolutely right. Okay. So you, so you really don't see Yarn really focusing on these, what our legacy or whatever infrastructure purchase has been made. It's moving pretty much for Amazon or Google or Azure. You, you basically just see, you know, the, all the, all the, it's not, it's not something to do with Yarn, right? I mean, Yarn added all the right features over the last couple of years in order to label a particular resources within the same cluster. So for example, you could today submit a job to, to a Yarn based scheduler saying that, Hey, I need to run this where I have two GPUs on each machine, right? As long as your infrastructure people have properly, you know, labeled those nodes with having those kinds of resources, it will schedule it there, right? But think about, you know, where most of the modern hardware is actually becoming available. Modern hardware is not becoming available on premises
Modern hardware is becoming available in the clouds first, right? So the new GPU that Nvidia comes up with or new FPGA that, you know, somebody produces, that's not coming to my data center first, that is going to be available in the cloud, right? And in that case, that supply side scheduling comes into picture, which is that, Hey, I don't have a fixed set of resources. Some of them are split into GPUs and CPUs, etc. I just basically go to my container engine. And I basically say, I need to create 20 containers in there that have access to two GPUs, right? And that becomes a supply side scheduling problem, right? So that's where most of the work actually, that's why we basically see that that that work is shifting there. The second thing is available of AI through API, right? You have SageMaker, you have MXNet, all these things, those are first getting deployed in production on the clouds, rather than making it on prem first, right? So all these modern use cases, they are going there first, because the modern API is for these machine learning workloads are becoming available first on the cloud, right? So that's why I think in the in the sort of next five years, you will see a lot of these workloads first getting onto the cloud, then second, the bill will come, and then they will get in migrated inside, right? So that's basically what is going to have to have. Yeah. All right, thanks. Thanks a lot.