Scaled: Matei Zaharia: Sparking the Data Revolution
Recording: Scaled: Matei Zaharia: Sparking the Data Revolution
Music Hello everybody, I'm Alexey Krabarov and here we are at an episode of the Scaled series, which is a podcast and video series and a blog about scaled companies, people and software. And here we are at Databricks on Location with Matei Zachariya, creator of Spark, chief technologist at Databricks, and a professor at Stanford and one of the co-founders of the Don Lab. So I think it's a really interesting episode for us. We're going to talk about Spark, which is an example of a system which scaled globally. And Matei was at the root of this, Databricks is a company which stands behind it, so I think it's one of the best examples of doing all these things. So thanks Matei for making the time. Okay, thanks a lot. So can you talk a little bit about the history of Spark, how it all started, give us a little bit of technical context, what made it a system which was interesting to you and then became interesting to so many other people
So, okay, sure. So Apache Spark is a large open source project now. It's got thousands of contributors and it's developed by a very wide community of multiple companies and individuals. But it started from this research project at UC Berkeley that I was doing during my PhD thesis, which was the Spark cluster computing engine. And what happened is when I started my PhD, this was back in 2007 that I started, I became interested in data center scale computing. It was just beginning to be, you know, a popular thing at the time. And I got to work, you know, I got to see a lot of the early uses of it through Apache Hadoop. And I got to work with a lot of the early users of Hadoop, such as Facebook and Yahoo, and see the kind of problems they were having
And in seeing that I saw, okay, there's a lot of potential in actually, you know, creating these large data sets and writing applications that scale out of our cluster. But there are a lot of applications that are difficult to do. And, you know, people could really benefit from a easier to use programming model and from something that makes it very modular, more like normal programming, where you can hook together a lot of libraries to get a task done and, you know, collaborate with lots of people to build that distributed software. So that's, that's the context where, you know, I started the work on the Spark Engine. Initially, we were targeting interactive queries using, you know, the Scala interpreter at the time. And we were also targeting iterative machine learning applications, because a lot of people in our lab, in the Rad Lab at Berkeley did machine learning. And over time, you know, we saw this is interesting. There were people in industry who wanted to use it
And there was also a group of other grad students who, you know, started working on the project and contributing to it. So we had a really great team from the beginning of multiple people who are working on, you know, different aspects of this system. And eventually, you know, as it grew, you know, we decided to, you know, move the project to the Apache Software Foundation and have this environment where, you know, many people can contribute. So it's written in Scala, which at the time was a new language, right? And a few people knew about it. What made you write it in Scala? What was interesting in Scala for you? Yeah, that's a good question. So, so first of all, you know, when we started, we had no idea that this would necessarily become something that people use, right? Because the goal of research, academic research is usually to build prototypes of things to show some idea. It's not necessarily to build, you know, production software. But so, so Scala was, you know, I wanted to learn Scala actually, and I figured this is good
You know, this is a good project to do it on. But also, I really liked the support for functional programming in Scala to have a very concise programming interface. It seemed really clear that people in practice want to chain together many distributed operations like, you know, tens, maybe hundreds of MapReduce steps. And it doesn't make sense to ask them to write, you know, a whole separate Java class for each one and so on. It's just going to slow them down significantly. So, I really liked the functional programming syntax. And there are a couple of other reasons that we did in Scala. The reason I even heard of Scala in the first place was actually Michael Armburst, who is also a PMC member on Apache Spark
And he started the Spark SQL project. He started Delta Lake after. He's also like a, you know, senior staff engineer at Databricks. So, he was always, in the lab, he was always, you know, the PhD student who was ahead of the times in terms of technology. He was the first guy to get an iPhone. When the iPhone came out, he was the first person to use Git. Okay. And he was also, you know, he learned about Scala and he was using Scala in one of his projects
So, I learned about it there and thought, oh, this is pretty cool. Let's try combining it with MapReduce. And there was also a project by another grad student, David Hall, who had built a Scala API for Hadoop, who showed that, you know, you can get Scala to compile into these classes and then ship them over the network. And so, I saw that it is possible to do distributed computing with it. So, that's how it started. It wasn't, you know, necessarily that we said, okay, this is the best language, like, for production in this space. But it turned out to be great because it did build on the very strong Java ecosystem. And it made it a lot better for this kind of exploratory data analysis and even for production, you know, data analysis jobs
It's just use the right tool for the job. Right. And then David Hall then produced a lot of very interesting Scala Soterfer NLP. Right. Yeah. So, yeah, it's a small world because we see these folks, you know, in the community doing great things. Yeah, exactly. So, I'm very curious about the RDDs, right, because the RDD, I think when Spark just came out, it was an RDD paper
And so, RDD was a very interesting concept. Can you talk a little bit about them and, you know, how you came to this idea? Yeah. So, what was happening at the time is, you know, everyone was interested, especially in the tech world, in these large-scale computing frameworks. And people were designing many custom, you know, distributed systems to execute different kinds of computation in parallel. So, at Google, for example, there was MapReduce. They wrote about how they used that to build, you know, their web index and so on. But then there was a separate system for SQL. There was Dremel and then later F1 for SQL at scale
There was the system for Graph Computation. You know, there were a number of other systems. And the same thing happened in the open source world. You had all these kind of custom systems built for things like streaming or machine learning or graph processing and so on. So, we realized pretty soon that, you know, even though it's nice to do all these computations, just managing and using one of these systems is really hard by itself. And then if you ask someone, well, you know, just install and administer, you know, five or ten of these and hook them together, that's going to be really hard for them to adopt. And they also, they each had like slightly different data formats and APIs and so on. So, we, you know, we quickly kind of decided to look in another direction, which is can we have, you know, one programming abstraction that actually covers all these use cases efficiently
And the one we came up with was RDDs, which is basically is distributed collections and the initial version, they would hold just objects like Java, Scala and Python objects. And you could control the partitioning across nodes. So, you could control what data is co-located. And that's really important because co-locating stuff is basically one of the main performance optimizations you can make in these systems. That's one of the things that was different in each of the systems is how they partition data. And then you just had a kind of this menu of operations you can do on them like map, group by and so on, you know, zip, stuff like that. And those were also aware of partitioning and they did kind of the obvious optimization. If something's already partitioned by key, you can group it locally and stuff like that
So, that was the idea. And then we showed that, you know, you can actually express a lot of computations in it. You can still do the very fast, you know, single machine stuff that everyone was doing for their custom operation. And then you can hook them together into a fault tolerant distributed job. So, that was the, you know, the idea with RDDs. Which is resilient distributed datasets. Yes. Right
Yeah. This struck me really that, you know, you described how if a machine fails, the RDD can be recomputed. Yeah. And to me it was very interesting, right, that basically I think they wrote this notion that it's, you know, declarative computation which now can be recreated, right, if something fails. How, I wonder, right, like you guys are at the research lab, right? This is not what typically academic researchers think of, you know, being resilient and fault tolerant. So, how, like, this idea came about? What was influencing you to think forward to basically in production terms? Yeah. So, in this case it was very clear that, you know, the industry cared about resilience and fault tolerance. And there's also a tradition in distributed systems research, at least, to look at these
You know, not in all kinds of research, but definitely there are some areas that are concerned with it. So, we definitely, you know, we felt like we had to do that well. And it also leads to interesting problems to think about, you know, theoretical problems or design problems, which is always good in research. You want to have, you know, some, like, some hard problem that comes out that you can think about. So, it was just kind of fun to think about. Mm-hmm. Yeah, you know, so, in this series, we are looking at things which separate the things which scale from those which don't, didn't, right? And so, obviously, many stars aligned for Spark. But, so, one thing, as you mentioned, strikes me that when I played with Spark, it was first at a company called Cloud, where essentially, we had a data center, but we also have Amazon
And Spark was available at Amazon, right? And so, I loaded it using the scripts, which I think were in Python, and it spins up a four-node cluster, right? We're talking about 2012. And I uploaded a bunch of data to S3. So, you know, the first time we met was when I actually tried to load data from S3 into Spark. Mm-hmm. And it didn't work. And then I asked, can you go, she's like, can you help me? And, you know, we met in a coffee shop in Berkeley. And I think the reason was the timeouts on S3. It was set differently from probably for the better access
You had probably faster or something. All right, my timeouts were too short, so I could not load it. And so, that was the initial setup. And so, what struck me once we got this working, essentially, the cluster became my personal computer. Mm-hmm. Because now we have REPL, which is in memory, you know, prompt, basically interactive prompt, and all the data is in memory. So, suddenly, I can grab, basically, a bunch of data, right? Like, I can iterate. And so, in Scala, because everything is a collection, right? I could have a collection
I could just say dot map, right? And then I would provide a function. And so, what struck me is, essentially, the whole big data could be reduced to map filter reduce effectively, right? So, because any function you can think about, you just say dot filter. It will start, you know, tap complete, and it will be there. So, what, for me, what really made it click was that, basically, if you're a functional programmer, you already think in terms of collections, right? You already think that data is a collection, which is very different from the way Fortran was doing with do loops and kind of Python with, even with loops, comprehensions, and other languages you did not think, and in Java, you did not think of things as collections, right? And Hadoop was made this way, but it didn't, it wasn't clear, like in global Java. So, Scala made you think of the collection, suddenly, Spark just makes this collection. It lifts it on the cluster, and all the things you can think about are already there, right? So, I wonder, right, if this is something which you found clicked with other people, right? Because as a programmer, immediately, it resonates with me. I wonder if that was one of the things which led to widespread adoption in your mind. Yeah, I think this definitely helps
So, there was this concept of functional programming on collections that people could already read about. There were well understood benefits, and there were people already doing it for various purposes. And it's also a really good fit for distributed and fault tolerant computing, because you want things to be replayable, and you probably want things to be lazy, so that you have time to optimize the computation once you see everything the person wants to do. So, definitely clicked with, you know, with a lot of software engineers. And you had to instrument REPL early on, right? Because you needed to ship continuations to other nodes. Can you talk a little bit about that? You know, because, you know, basically, from somebody who learns the language, you now become a hacker. So, how hard was that, and what were the main challenges in the initial design? Yeah, that was pretty fun, yeah. So, the first thing we did with, you know, in whatever the project that became Spark was, we just supported compiled applications
So, I saw that when you compile an application in Scala, and it uses closures, each closure becomes a Java class, you know, that has a certain API. So, I could just take, you know, like, look at the class that they sent us, or serialize that object, and then call it on a different node. But I needed to have that class file accessible on all the nodes. So, that was fine for a compiled application where I have all the code in advance. And then I wanted to say, hey, you know, I just, I had the idea that it would be really interesting to do this interactively. Like, that would be a killer demo if I could just do this stuff interactively. And I looked at what does the Scala compiler do to actually compile stuff. And because the Scala compiler ultimately has to deal with the JVM, which has to load classes from somewhere, you know, I saw that it generated, I think it generated these in-memory, you know, bytecode objects for the classes
And then it pointed the Java class loader to them to load them. So, a lot of this is possible because Java has dynamic class loading, which is interestingly because Java was designed for loading code over the internet and applets and stuff like that. That's why it has it. So, I did that. And so, I modified the interpreter without really knowing exactly how it works to put the classes into a directory on disk. And then I ran an HTTP server on that. And I knew that in Java there's something called URL class loader that lets you load classes over the internet. Again, for applets, which no one uses anymore
But it's one of the thing, if you learn Java, you know, when I was learning it, when applets were all the rage, this was one of the first things you saw. It's like, oh, it's so magic. There's URL class loader. You can just load your code. So, I just set that up on the workers. And, you know, once I got the class path correct and got that loader into there, it actually worked. So, there were a lot of corner cases and things. If you look at how we do it now, we do some additional work to make sure that the closures can be sent
And we do a bunch of work to prevent users from shooting themselves in the foot. But that was the idea. And then once I saw that, you know, it made for a really cool demo. And I got really excited about interactive big data analysis, which you couldn't really do in a programming language at that point. You could, you know, maybe you could do it in SQL, but certainly not with custom code. Right. No general purpose language. Yeah
And then I remember vividly that, you know, that changed basically everything. Because at that point, Cloud was using Hive. Right. And so, Hive basically, this kind of SQL implementation on top of HDFS. And you basically wait sometimes minutes for a query. Right. And so, that basically killed the idea of ADA. And suddenly with Spark, now you just load
I think your initial demo was the Wikipedia. You can suddenly load the Wikipedia and then you can grab how many mentions of Stanford. Are there versus the mentions of Berkeley? Yeah. I don't remember who was, which was more frequent. But basically... It was a head, yeah. Right. I have to admit, yeah
But like, I mean, that, like in 2012, it was unfathomable. You have the whole Wikipedia and you just grab it for answers like this. Right. So, I mean, to me, the potential for EDA was very, very obvious. Right. So, now let's kind of look at, so I think technology is really scalable. And so, I think in this series, we really want to get to kind of the roots of scalability through kind of advanced technology, but also how community kind of takes it and how companies take it. So, if you look at the community, I think Amplab did something really new with the community
Because I remember you started these camps, right? And the camps were really the seed of the community uptake. Can you talk a little bit about how this camp started and what kind of happened in the camps and how that helped Spark adoption? Yeah, we did a few different things. And honestly, it's hard to tell going back what had the most impact. So, one thing at the very beginning, we actually spent time with those brave early users who actually, you know, wanted to try out the system. You know, we visited them. I remember I used to know, you know, all the users and I would see them every couple of weeks and so on. And eventually, I started running into people who were using Spark successfully without me knowing them. So, that was really great
But that's at the beginning. Another step that helped was meetups. So, I think you suggested when we met, at some point, once we got to know each other, you suggested, why don't you start a meetup on this Spark thing? And Cloud actually hosted the first meetup. So, that was great for bringing together these users and letting them share stuff. And then, yeah, at the lab, at that time they had started Amplab, a new lab focused on big data. Specifically, and we decided to host these camps. And I think for the first camp, Mike Franklin, the director, only approved the idea because we had a funny name. It was called AmpCamp
Yes. So, initially when we, you know, pitched it to him, he wasn't that excited. But then we said the name and he said, okay, let's do it. Right. And that let us do this short tutorial for people and record it and get people started. But I think there's a little more to it than that. So, the way I think about it, any software and really any product or whatever, you know, it needs a good amount of use to make like the critical, you know, journeys through that product work really well. And so, you need some critical mass of people that will use it and will give feedback and will let you smoothen those and make it really good for at least some use cases
So, at the beginning, I think we would have had to do a lot of stuff manually to make it good. We would have to listen to feedback. And then it became good enough for some things that people could do that on their own and could push it in new directions. So, I think that's an important part of it is think about, you know, what is it supposed to do for people? Is it accomplishing that? You know, why not? And make that smoother. And once we did that, like the ideal thing is if it's so easy to use for at least some use case that people can get started without any help, without coming to a two-day camp or whatever. So, that's I think what helped that go a lot after. And I remember when, you know, you guys did tutorials, I think, at Strata and they were really well done, right? Because basically everybody would get a login and we get a cluster, right? And essentially it was not a toy application because you basically, I think you had some credits on AWS. Yeah, yeah
Right. And so, how did that idea came to basically give people like this real world ability and obviously it was a lot of work on your part. I mean, I wonder how did that idea occur and how, what did you learn from this kind of tutorials? Yeah, that's a good question. Yeah. So, ease of getting started both when you download the software and also when you do a tutorial was super important. And fortunately coming from this, this like university background, we had a lot of people who are good teachers and who had taught classes and design labs and stuff like that. And everyone was passionate about it and, you know, everyone had opinions on what would make it good or bad. So, we had people who had to, you know, write scripts to like, you know, enable a class of a few hundred Berkeley students to do something
So, they weren't afraid of doing it to set up clusters on EC2 and get people to try Spark. We had people who taught, you know, like intro to machine learning or whatever. So, they could teach people that. So, I think that helped. And we, you know, we knew that if people have a great experience, they're more excited to try this out later. And we had a lot of fun doing it. So, we developed all those materials over time. They became a lot better over time
And at that point, you know, when we saw people actually using the software for, you know, real use cases, we also got really excited about that. And we just, we wanted to do it for that reason. And we also knew it would let us, even from the research perspective, it would let us do much better research because we can understand, you know, the applications and the things people want to do. But, yeah, it started, it took a while to iterate on that. But we set pretty high goals from the beginning. We were pretty ambitious about what can someone do within a few hours of trying this out. Yeah, I remember vividly, you know, the early days of Spark. I think it was really, you know, super popular because, you know, the rooms that you guys talked at big data conferences were full, right? Like, they basically have to kind of, you know, push back on and the tutorials are full, right? So, I think that was really, I think there was a confidence of time when there was this pent up demand
And you really met it well. And so, in terms of application, I think it's very interesting because obviously Spark is perfect for ETL, right? Like, I found that essentially anybody who had giant amount of JSON or, right, like text, right, HDFS, whatever, like it was immediate replacement for Hadoop because you now can take all this data, weblog data, right, server logs and basically extract it. And so, I think the challenge, I think it's still present, right? That it was on JVM. So, ETL is primary driver. Data science and machine learning mostly happens in Python. I think it was not as pronounced in 2012, right? But then we basically saw huge influx of kind of the boom of data science happened after that in machine learning and now AI, right? And it mostly happens in native. So, and Spark made a lot of progress in interplay, but can you talk a little bit about, right, so the use cases because now we have ETL which we know how to do and machine learning is, in data science, it's also diverging. People have been called different things
They called analytics, right? And obviously, systems evolved after that, like TensorFlow and others. So, how did you see the evolution? So, here's the undeniable, super easy ETL use case, which we all know how to do. And then you guys went into different directions because there were data tables and now MLflow and all that. Can you talk a little bit about this now evolution, how it was informed by all these use cases? Yeah, yeah. So, yeah, so definitely Apache Spark evolved over time in various ways. And also, you know, we wanted to support new use cases that were coming out such as data science. We actually had, you know, machine learning was one of the initial driving applications. That's the thing that people couldn't really do efficiently with map pages before that we targeted
But, yeah, so a few things happen. So, in general, I mean, our goal is just to make, you know, parallel data processing easier, whatever purpose you're doing it for. And most data processing, you know, you have to chain together different types of operations, different types of computation, and you also have to go in and explore what's happening, right? When your pipeline breaks, you want to go in and see, hey, what was happening with these records or whatever. So it makes a lot of sense to have one API that, you know, can invoke all those computations and do things with them. So there were a couple of things that happened. So one thing was definitely the increased focus on data science and Python. So today, I don't know about worldwide among Spark users, but definitely among users of Databricks, you know, the platform where we can see what people are doing, Python is definitely the most used programming language. And there's also a lot of SQL, actually, where, you know, people just want to quickly look at some data and see some results, and they'll do that in SQL
So we wanted to make sure that we support those really well. And I think to some extent also, like, you know, adding support for Python and adding support, like, for example, we use Apache Arrow to make the data transfer faster. And we have this API Spark data frames. And now there's this other one, koalas, that's very similar to pandas. This also made Python better for large scale stuff. And it made the data scientists happy in their organization that they could actually scale up the Python to a, you know, large data set and not have to, you know, get some engineer to rewrite it in a different language. So that's one thing that happened. And then another thing that happened is for machine learning, we wanted to make it really easy to use, to combine Spark with any machine learning library out there
So today there are really good connectors that let you run things like XGBoost and LightGPM, H2O, TensorFlow, Horovod, all within a Spark job and feed data between them. And that's really kind of the ideal for us because there's so much innovation in the area. And we just want to make sure you can combine and coordinate, you know, best in class sort of libraries and algorithms. Same thing as you could, you know, with just ETL, you can use Spark and call any Java library or any Python library at scale. So that's, you know, our high level goal is let you build and coordinate these massively parallel workflows. And it turns out by having good connectors to a lot of, you know, other libraries and a lot of data sources, you know, we can easily make that happen. So, yeah, so now we're going basically into the world and right then, Databricks is a hosted Spark, right? And so you, I think when I was at Nitra, we were one of the early users of Databricks. And it was, it was very nice and kind of, you know, to see it evolve right into, I think, really a huge, huge company
So I wonder, right, what from your insights, what happens here? Because I think we kind of see now the magic of scale because here we started with a bunch of graduate students who are some of the best in the world and they're in this lab and magic happens. So the scale is born because you have the hunches, you have the intuition what's going to work, right? So you hit upon this project, it was a success. You really kind of shepherded it with the community for the back, right? And so, and it's scaled as open source project, it's scaled as a community. Now we have a commercial company and you go into the world and in the way you have to repeat the scale, the magic of scale, right? Now at the business level, it's a different game, right? And now you basically have a whole bunch of customers who are not technology first companies, right? So you have, you know, kind of media companies, you have hotel companies, you have companies whose businesses it is to run something other than a bunch of computers, right? But they use data. So I wonder, how do you take this kind of magic of scale and this kind of intuition that stuff which really works, how do you convince other people that this is the way to do it? And how do you see them scale? I wonder if you can kind of generalize it a little bit from the customer experiences, what really works in the real world, you know, based on open source software? Yeah. So, I mean, there are a lot of different things to say there, but let me, I guess, let me start with, you know, what problem we're solving for people. And then I can talk about, you know, there's also more tactical stuff you have to do in growing a company and, you know, actually finding customers and making sure the product is great for them and stuff like that. But the core thing that we're trying to solve in both Apache Spark and in Databricks is we just see there's a lot of potential for applications powered by data, especially large scale data or new algorithms such as machine learning and they're really complex to build
And they don't need, there's no fundamental reason why they have to be that complex. So we think we can make it simpler. So what was happening in, let's say in the initial, you know, Spark project at Berkeley, for example, is we saw, you know, people are trying to run these distributed jobs. They have, they just want to apply a function to all the records, like a filter or a map, but they have to jump through all these hoops to make it happen from a, from a programming standpoint. So we tried to make that easier for software engineers. And we came up with ways, both like a design of an API and also design of an engine that can hide a lot of the complexity from you. Then same thing with, you know, with data scientists, data scientists know statistics. They know, you know, various transformations on data
But as soon as it becomes larger than, you know, what fits on their laptop, they have to deal with all these other concerns. So how can we get rid of that? So we got things like PySpark and then the Spark data frames and Koalas and Spark MLlib and stuff like that, lets them, you know, lets them do that. And there's no question that now they can, you know, they have to worry about the distributed stuff and the cluster setup and all that less often. So it's made it easier for them. And I would say the same thing is, is true with, with Databricks. We see a lot of organizations who have large data sets. They have talent there that knows statistics or knows, you know, science relevant to their business or something like that. But they don't have the, you know, the talent or the resources or the time to also set up a giant, you know, software engineering and DevOps organization
To make that happen. And so one of the key things that we provide because through being cloud service is we can do that for you. So you just focus on the logic that's specific to your data science or statistics or whatever problem. We'll make sure it runs reliably. We'll make sure it's secure. We'll make sure it's performant and so on. And that just really lowers the barrier for many of these applications to happen. You know, there are tens of applications where they wouldn't, it wouldn't justify hiring, you know, and putting together a team of like 10 people just to get that expertise
But it would justify, you know, using this service and writing some Python or R or even some, you know, Scala and getting this to happen. One way I like to think about it just to set up an analogy is if you think of the early uses of computing and business in general or even of something like web applications. So if you remember, it used to be to build an interactive web application that was really rocket science. You, you needed, you know, you needed like Apache web server, you know, CGI, various databases, forms. Yeah, stuff like that. And, you know, very few companies had, you know, interactive web applications was, you know, it was a big deal if you actually had an online store or, you know, even like a website that did something really interesting. But it's ultimately is the same problem that everyone is trying to solve. And we develop technology that makes it a lot easier
And today companies have not just, you know, one web application, like, you know, the thing to book your flight online or to view your bank account balance, but they have thousands internal apps, external apps. And these are all apps where you wouldn't want to spend, you know, two years building each one as you had to do back in the day. So we discovered, you know, technology to wrap up the common patterns, make it easier. So that's what we're trying to do with data as well. So first through individual libraries and operators you can chain together in Apache Spark, and then through the SaaS and production aspect of Databricks that lets you actually run this thing reliably, you know, all the time. So I think there's a lot of potential here. And that's, this is one of the main reasons I think that this stuff has killed is because people who couldn't justify using the technology before it was too much of a hassle can now use it. And there's just like way more of those than there are, you know, of the people who would have used that before
So every year there's new workloads, new companies that use it. And, you know, just, I mean, just to give some example, like a lot of the companies we work with, it is sort of their core business to do something with data. For example, think of an insurance company. The whole point of insurance companies is to, you know, to predict risk and somehow win out when they set their payment and they compete. And they've been doing that for hundreds of years and they're successful. But when they want to run like, hey, let's, can we run this at a hundred times larger scale? Or can we bring this public data set that's like, you know, half a petabyte or whatever? We want to enable those same people who build those models to reliably run at that scale. So that's, these are the kinds of applications that we get. And I remember you gave a talk recently at an ICM conference that you guys move in petabytes of data
Right. And so you managed to build this infrastructure. So actually, that's a very interesting point. I'd like to touch upon and see your insights. So I think the initial users of Spark were software engineers. Right. And so they knew things like operations, cluster management, if something was wrong on EC2, you know, they could associate into machines and look at them and do stuff. Right
And so I think, and especially on people who run, you know, things on-prem. So, and then the users, you know, kind of more senior software engineers, so had experience for new functional programming. They obviously knew Git and versioning. And so that stuff was kind of, that's something that they knew. Right. Now we move into the data science and we see a lot of folks who learned programming in Python when learning data science. We're going through some bootcamps or programs and, or, you know, applied statisticians, mathematicians. So we have a lot of folks who did not have this traditional software engineering training
Right. And so they are obviously a huge market. I think this is driving the growth of Python and data in general, because that's the first easiest kind of step they can make. Right. On the other hand, they do not have the experience of software engineering. So we have a lot of concepts which we take for granted in software engineering moving into data science. So I think that this whole model deployment market was a bit surprising to me at first, because suddenly a bunch of startups proliferated that do model deployment. And I wonder why is that, right? Model is just an object
You serialize it, right? You deploy it by using whatever you do to deploy serialized blobs of data, and you version it using all the things you know. You make checksum it, you mandate it, you can put some hashtags in Git, if not the whole thing. Like what's so hard about it? And suddenly we have this whole market. And I think it actually evolved that where not only models deployed somewhere and rolled back, but it also can be automatically tuned using some hyper parameters and monitors and so forth. So I wonder, you know, how you guys deal with this, because obviously you have a lot of folks who started software engineers. Now you encounter a bunch of people who are not. And I think we have this kind of three legs. So we have software engineers, data scientists, and business owners
And so the model deployment, I think, is successful, this whole business, because it touches upon this whole three things. And so you come in as a communicator, right? And you mediate these three things. And it's the first time in business when business owners now basically see models as objects making money or not. And so now they become involved. That's my read on this. I don't know if you see that, right? But I wonder what do you guys see and how do you see this interplay of software engineering skills, data science? Can you take it all the way from data scientists and manage it in the cloud? Or is it some skills which we need to teach data scientists? Do we need to bring it into functional programming, for instance? And maybe should we consider Python as an intermediate stage and then, you know, going somewhere next, which is type safe and so forth? So how do you see this interplay of these components? Yeah, it's a good question. And I think there's no right answer for everyone. So I think you have to think about kind of the best approach, the best team setup for each task
So as I said, one of our big goals is just to simplify working with data for everyone. So that naturally means make it simpler for non-software engineers as well as software engineers. And because everything is based on APIs, it's possible to do both, right? You could write code that talks to Spark or Databricks in a notebook and not store it anywhere. Or you could write it in a, you know, in a file that's in a version control system and answer your test infrastructure and all that stuff. And you both benefit from the same thing. So that's good. But I think you need to think about the use cases too. So if you think about something like exploratory data analysis, you want it to be super quick, very lightweight
You don't want to have to set up a Git repo and, you know, like a Kubernetes, you know, spec and all that stuff just to do some exploratory. Data science. If you think about, hey, here's some job that will do, you know, billing and production for us and needs to not fail. That's when, you know, you do want to set up all the software engineering stuff. And we see even software engineers, even teams, you know, with a lot of software engineers, they don't want to spend their time, you know, heavily engineering each thing. If there's some insight they can get through quick exploration in SQL or Python, that's awesome. They'll do that. If they can, you know, press a button and turn that into like a dashboard that gets updated every night and they can post it on the wall next to their team, they'll definitely do that
And then when there's something that's worth, you know, spending more engineering time, they'll do that because they're busy, you know, can't spend time on everything. You have to decide, like, what do you spend a lot of time on versus a little bit of time on. So I would say it depends. I think a lot of data scientists are learning software engineering concepts and tools, and that's great. But there are also a lot of people, including people who know those, who just want something simpler. They just want to go in and within, you know, two minutes, like figure out, you know, whether their feature is being used in the product or whatever. And for them, you know, you can also provide some really good tools. And again, the end state, you know, if you think about it, like if you think of my web application analogy, there are a lot of things where, you know, you don't spend too much effort on it
You can get a basic application, like let's say a careers page for your company or something like that, right? How many people build their own? And then there are others where you engineer every aspect of it yourself because it's a custom thing or it's a highly differentiating thing. Right. And you want to be able to do, you know, either one. Right. You want to continue a focus. So I want to kind of shift gears a bit and look at retrospective, right? So we have a really successful system in Spark. And what's interesting to me, there's not really many equivalents. So Spark is written in Scala and obviously PySparks, of course, but what started as a demo for Mesos, I would say now Eclipse Mesos in many ways, because Mesos met a competitor in the face of Kubernetes
And the company behind Mesos had to evolve, right, and support multiple things. And so I think I talked to Ben Hidman in the previous episode. He, I think they learned a lot of things about this is the applications and they obviously some of the best people in the world who know about state and distributed systems. And so they can take, so they're now developing AI operator called CUDA, data science operator for Kubernetes. So they can take this, so there are certainly a set of transportable skills. So, but I wonder in your mind, like nothing like this emerged to compute the Spark. There is no distributed system in Google Spark in any other language. There is no system in Haskell, there is no system in Rust yet
And I don't know if it's going to be. So people are not trying to do that. On the other hand, Apache Kafka was written in Scala, right? So we have these two things. That's true, yeah. And I think both are really dramatically grown. So I call them actually, like my kind of way to reason about them, I call them second wave companies. Because you have Google and Amazon and Microsoft grew out of, even Google, you know, grew out of proprietary data centers and so forth. It did not grow as an open source company using this business
But Databricks and Confluent, I think they're very unique examples of companies which basically bootstrapped using open source. They're completely software phenomena, right? They do not, like you run on top of Amazon, you don't own hardware. And open source was used to scale and gain mind share. So, you know, commercialization of open source is hard. So there is a different question, you know, how it's going to happen now. But there is no denying that these two companies basically grew out of open source projects, distributed systems written in Scala, which basically caught like fire in the minds of developers, right? So what is it about, you know, the systems of Spark versus Mesos? did not, like it was a toy app for Mesos. Now it obviously, you know, generalize much further. What do you think distinguishes if there is anything, things like Kafka and Spark versus others? And why do we not see similar systems like that emerge? I think it's been a while, a few years
So this company has gained billion dollar regulations and like everybody's using Kafka, everybody's using Spark. What is this about? If you can, you know, any insights? I know it's a hard question, but, you know, maybe we don't know the whole kind of thing. But if you have to kind of identify what makes these projects different from other projects. Yeah. So I think it's hard to tell exactly in all cases, but I can tell you a few different things. So, so one of them is just, you know, what, what problem does it solve? And, you know, how important is that problem? How many people need it? So, for example, like Mesos is awesome technology, but there are fewer people who have to operate the data center than people who want to run a distributed computation, especially with the public cloud, because you can use the cloud to run it. Then you don't need to manage your own machines. So that's, that's one factor
Another factor that I think is, is, is really important is have the projects evolved over time and how, how nimble are they? And like, how, how can they keep up with things? So if you look at Apache Spark today, for example, the majority of usage is through data frames and the SQL optimizer, which is actually, it's a lot easier for end users than, than even the functional programming API was, especially, you know, for people who are using data frames. You know, for people who aren't software engineers and thinking about like, you know, memory use and stuff like that all the time. And it's a big evolution in terms of API, right? If, if we were writing the Spark research paper with, you know, today's version, we would have said instead, it's, you know, it's kind of like a database engine with, with very strong support for like user defined types and functions. Yes. It would be a, a different story. Same thing with Python, same thing with the integrations with deep learning libraries, right? Or when, when today, when people want to run deep learning, we say, Hey, yeah, it's got this really good data transfer to, to Horovat or to, to TensorFlow. So that's another factor is, you know, how, how well do they evolve? I think Kafka also has evolved in, in various ways. And then I think there's also in, in, in Apache Spark, at least there's really important effect of the ecosystem, basically a network effect
So, you know, people, everyone wants to run different types of distributed algorithms and access different types of distributed data sources and so on. And in Spark, it's easy to create connectors and it's easy to package up any algorithm, even something written in C++ or in Fortran, as a function that you can then run on an RDD or on a data frame. So it just, just because there are all these other things that work in there, like hundreds of packages or built in libraries, it's natural for people to try to, to build on top of it if they can, unless there's some hurdle that stops it. And, you know, we're always looking at what, what's hard to use about it? How can we make it easier to make sure that we don't have those, those hurdles? But it's kind of, I think an, an example of this is if you think about, let's say open source, you know, operating system kernels, there's only one really that's widely used. It's Linux, but it's highly flexible. It works on everything from like, you know, a little like embedded device to, you know, a supercomputer. And it's, it's got this big ecosystem around it. And as long as people can extend it and use it for what they need, that's, I think that's just kind of an equilibrium that can happen
Of course, there's always, you know, I also go back to what I said before, which is, I think all this, you know, data and machine learning technology is still harder to use than it needs to be. So I think there will be, you know, things that make it even easier. And we're hoping to, you know, add some of those in Spark or around it, things like Delta Lake and MLflow and so on. So I think that's, that's, that's why that's at least one of the benefits of using it is you get this community and this network of packages and things that all work together. Yes, absolutely. And I mean, and I've seen Spark evolve, right? And I think it's really interesting how you guys really grow with customers, right? And, and, and use that as a, as a kind of this feedback engine, right? To, to develop. So I wonder if we can kind of bring it to today. If, so what are the challenges, what opportunities? So I've see, we've seen talks about wealth, which is a project coming from, from Don and we had a great presentation at Scale by the Bay
And I think really folks see, you know, how it evolves. Can you talk a little bit about how this native execution, the rise of Python, which is native, how it, you know, kind of influences the development of data science, how it developed, influence development of Spark and why, like what motivated wealth for you and maybe Ray at Amplab, how these things are responses to what's happening in, in data science. Yeah, I think there are a lot of interesting trends actually, in terms of what's challenging for people to do. So, Weld targets kind of the performance fund to make it easier to get high performance underneath these data analytics libraries. So if you're not familiar with Weld, I guess it's, it's basically like a compiler and intermediate representation for data parallel code. And then we have these bindings that, that you take, you know, applications written in NumPy or Pandas or Spark or other libraries and, and compile them down to efficient machine code. So that project was actually motivated by, you know, our experience building the catalyst and the Spark SQL execution engine, which compiles, you know, applications to Java bytecode. I was wondering, can we make that easier? Cause it's a lot of work
And it's also inspired in part by, uh, sparks, um, like I think what, one of the main things that made Spark successful is that you can efficiently compose lots of operators or functions or libraries that were written even by completely different people. So for example, I can have a, a data source, let's say for reading from, uh, Cassandra and then I can have, um, you know, machine learning algorithm upfront that, and both of them talk through the, uh, data frame API. And because we see that the machine learning algorithm only reads a few columns, we can tell Cassandra to just read those columns or to apply a filter at the source or something like that. That's a really powerful, like cross, you know, function optimization. Um, so that's what inspired weld and, uh, that's, we're, we're definitely looking at improving, uh, you know, sparks efficiency, both in terms of CPUs and also in terms of like emerging accelerators. For example, um, uh, there's been a lot of work in the, in the open source project from Nvidia, from the rapid steam, uh, to enable, you know, spark to work with GPUs. So that's one area. Um, but there are also other areas
So for example, a big one for us is, um, is actually data management and transactions and indexing on these really large data lakes, like think, you know, billions of parquet files and so on. So when we started, uh, Databricks, we thought, okay, Amazon S3 or like Azure Blob Storage, these are really awesome storage systems. They, they've solved the storage problem. They're reliable, they're large scale, you know, they're not going to crash. It's good. Uh, so we're just going to do computing. Um, but then, you know, as people started using it, about half of the, you know, support tickets we saw coming in were related to storage. There were things like I was, um, running a job and then it crashed partway through and now my table is corrupt
Cause like some of the tasks sort output and some didn't, or there were two jobs trying to add to this partition and you know, they overwrote some stuff or there were eventual consistency things. I wrote some stuff and then I tried to read it and the counts were different. Um, so that, that was a big issue. Um, and even, uh, within Databricks for our, our data pipeline, we had this issue. Um, and so Michael Armbris started this, this project Delta Lake, which is a transactional storage layer on top of, uh, S3 and Azure storage, uh, that adds, uh, both acid transactions. So you don't have these reliability problems and more efficient, uh, indexing and metadata management. So you can quickly find relevant files for Quay. Uh, and, uh, that made a huge difference
It dropped all those storage issues to a small fraction of what they were before. Uh, and it enabled a lot more people like this was the biggest obstacle for a lot of teams using Databricks and, and Spark on it. So it enabled them, uh, to, to get past that and to reliably work with this stuff. Um, and, um, you know, so that's, that's just an example. And on the machine learning side, we also see a lot of issues with productionizing machine learning, monitoring it and so on. So that's why, uh, we have this project MLflow. So we're looking pretty broadly at what are the top issues. The nice thing about doing this at a company especially is that, you know, you, you have demanding customers, you can talk to them
And if you see an issue, you can quickly validate it with them. Um, so we're taking a pretty wide view. I would say like, um, Delta Lake, for example, is, is probably the most impactful out of the three things you mentioned, but compilation and machine learning, um, lifecycle are, are both important as well. You know, this strikes me, right? Like history repeats itself because, you know, my challenges with Spark started with those three. Yeah. Right. And so, and so, and so, right. And this is, this comes back because, and, and what, what do you describe effectively Databricks rebuilds the storage layer
It rebuilds SQL. So basically you guys come back to you, you rebuild the database for the new age. Yeah. Right now it's a database using public cloud for storage, right? And public open API. So, so that's very interesting to me that, you know, probably SQL is going to outlive all of us because it just keeps, keeps coming back. Well, SQL, yeah, SQL is very powerful. And there's also, um, really, it's like really well known how to build, um, you know, like good engines for it. So, uh, there's a lot of expertise in how to do it, a lot of examples to look at, but I think there are also some differences
Like I think the whole big data ecosystem has basically created an open or modular type of database engine, which just didn't exist before. it was all the vendors, they said, you know, you don't touch our internals. We do everything. Yes. You, you just write a SQL and that's enabled people to do these really powerful applications that require, you know, code as well as SQL. Yes. Um, I also think maybe the future, you know, like data warehouse or whatever is, is going to involve, is going to have models. Machine learning models as a first class data type, maybe machine learning features
Um, definitely streams are, are something we're seeing a lot of. So, so there are some, um, you know, some, some, some differences, um, uh, as well. And datasets maybe, right? I think versioning models and datasets. Versioning is a big problem. Versioning is super important. Yeah. Um, so that's, that's another thing where we can kind of redo it maybe in a, in an easier way. Um, yeah
So, uh, another topic, uh, so we're coming to an end. Uh, I wonder, like if you look at the future, you know, some of the themes which you mentioned, it strikes me that, you know, you almost talked about serverless because you said, you know, Spark enable people to run these functions. Mm-hmm. Basically, it's where lambda functions. When you're shipping the closures, you're shipping functions. Mm-hmm. So we have, in a way, we're talking about serverless in some way because now you're, we can talk about functions and it's in functional programming. We all know that basically, you know, you just want to run lambda functions on, on things, right? Okay
And that's how we think about it. This is the way functional programmers think. This is not necessarily the way, uh, data scientists think immediately, but I think it should really be appealing to everybody, right? So, Spark is in a way, uh, I think it's kind of one of the earliest serverless systems because you have the function. Now, we don't care where it runs, right? You have the programmatic concepts running, of course, and now you're kind of removing the, the, the obstacles, and you have the public cloud. You basically want to let people run functions. So, in a way, if you're within Spark, I would say this is a serverless system because you run functions on things, right? Mm-hmm. But now we have the rise of serverless as a concept and people are kind of wrapping their minds and, at, uh, scale by the way, we had a panel, uh, a very interesting panel, uh, where basically we tried to predict in 10 years, will all the, you know, unicorns will be basically based on serverless computing. Mm-hmm
So, obviously, somebody will still have to run servers. And so, in a way, uh, then Databricks would be one of the providers of the serverless computations, because now you have, right, together with, with whatever public cloud is under there, you just get much, much easier contact. So, what's your take on the serverless? I mean, is this just a marketing term because we have everything already there? Mm-hmm. We just call it, is it just kind of ease of use term, which is a shorthand for a set of things, right? Like, in your ideal Spark system, you have a function and you can run it across different things. Mm-hmm. Is that that? Or is there something more, uh, what works, what is the hardest work we need to do to take advantage of serverless? What's your take on this? Mm-hmm. Yeah, that's a good question. So, I think the, from, from a user experience, as you said, someone using, uh, Apache Spark, uh, they are getting basically a serverless experience because they write, uh, an application
They don't need to know how many nodes there will be. Mm-hmm. They don't need to know if nodes come up or down, uh, and they just specify a function and then the engine will run it. Same thing with interactive use. Like, actually, when people use our clusters, you know, the, the cluster is actually auto scale up and down. So, like, if no one's using it, it's got very few nodes. As soon as you run, it spins up some nodes and, uh, it's the same, you know, sort of idea. We, we even keep, like, a pool of VMs that are hot that we can swap in to do this stuff
Um, but, so, so that's that part. Um, uh, and I, I think it is important. For, like, broad use of this that people don't have to worry, uh, about, uh, configuring machines and, and clusters and so on. Um, in terms of, um, if you view serverless, like, I view it maybe more as also a, a technology to have these micro containers on the back end in, in the cloud. And that is pretty interesting for quickly ramping up and down. Uh, I think it's the real, like, the biggest killer app there is if you just have, um, some application with pretty low, uh, um, request rates overall. Because you get something that's highly available, you know, maybe geo-distributed, but even if you're making, like, 10 requests per day, uh, you know, you don't have to pay for, like, those geo-distributed servers. So, I think that's very unique
That's what a lot of people use that for. Uh, for data analytics, it is interesting for scaling up and down quickly. And, um, but, but there are also a lot of limitations in the current platforms that often make it more expensive and, in, in, in both dollars and in, uh, time, compared to, uh, compared to just having, uh, you know, a bunch of VMs. So, I think we'll have to see how that plays out. We're definitely excited if, uh, if a new technology comes along or if, if, if this is open so you can write, you know, long running applications or have data caching or shuffle or stuff like that. Uh, it would be great to try to run, you know, for example, like, have a Spark backend that runs on top of this. Um, yeah. And, uh, I guess as, as programmers, we can always kind of apply, right? Because we know how to do with functions
So, I think, I think we're in a unique position. Yeah. It's great for users. Yeah. It's, it's definitely, and, uh, I could even say that maybe because of the success of functional programming in things like MapReduce, uh, that might be one of the things that motivated people to try serverless for, you know, micro services and stuff as well. It's just, it's just a function, right? Yes. Yeah. Yes
And so, I think the final question is, uh, I think you're in a very unique position, uh, you know, having been the person who created this at AMP lab as a graduate student, then you basically saw the OSS project with all the colleagues and community grow and then the company evolved and now you, you're, you're a professor at Don Lab. And, and so I wonder like, uh, what's most interesting direction for you because, you know, scale is magic, right? So do you, uh, want to see it work again? Do you want to kind of, do you see, well, there's a complimentary thing in Spark? Do you see new, completely new things of global scale coming out of Don Lab? Mm-hmm. And, uh, you know, what is the ingredients, right? Like you, you've seen it work in the template app, but I don't think, uh, like, you know, so like, I think from the vantage point of, of, of, from your vantage point, you, you, you see how Databricks works, what works in the world. And so you have this advantage of the creation of Don Lab that you have this ingredient of the same time you have some of the best minds in academia working on Don Lab. So I wonder, like, if you could come back to the magic of scale and if there is something new, you know, is to come from this mix, right? How do you see kind of what's most exciting for you to apply what you know at this point in your research? Are you motivated by technology, you know, such as Rust, you know, in building well, applying compiler technology to data science? And kind of how do you see that kind of, you know, what, what, what do you focus to kind of make the most impact? I would say it's an open-ended question. Mm-hmm. What's most interesting to you? How do you use technology to achieve it? Mm-hmm. Yeah, that's a great question
So, I mean, I think there's at least two areas. There's the, um, uh, kind of like, you know, more practical, like, um, uh, you know, immediate term type, uh, work that I, I'm doing with Databricks. And then research work is, you know, supposed to be long-term and high risk. So I'm also taking some bets, uh, there and doing some things there. Um, so first of all, I'm, I am really excited about, uh, the cloud and services like Databricks as, um, as really democratizing computing and enabling many more people to build highly reliable, uh, you know, scalable, uh, production applications. And with Databricks, we're looking, you know, we're doing this with data and machine learning in particular. So, um, I'm, I'm really glad that we chose the cloud model. It's a, it's a great way to, um, uh, to basically, you know, compared to shipping someone a bunch of bits, the cloud service can provide a lot more value because we manage it
And it means a lot more people can use it that, you know, couldn't have used the other stuff. Uh, so it has impact and it's also a great way to give us really fast feedback and keep us honest and, you know, not, you know, we don't have to, uh, build something for a year and then wait another year for someone to deploy it. We can just see immediately whether it worked. If we're wrong about some idea, we can say, okay, we were wrong about that. We'll try something else. Um, so, so I am excited about that. And I think there's a lot that can be, um, you know, that can be done to make, uh, data and, and machine learning, you know, applications, uh, easier to build. And, you know, I'm really excited to see like some of the things that people are doing, whether they're, you know, software engineers that are, you know, just now they can focus on, you know, their logic and get something reliable or data scientists or statistician or like, you know, bioinformatics people who can suddenly access these large data scientists
that they couldn't before. Um, so that's one area. I think there's always new issues coming up and I, I'm also really, uh, passionate in there, not just about APIs, but also, uh, things like user interfaces and just the process by which people work and how they discover that, you know, whether something is working or not. Um, on the, on the research side, uh, the, the Dawn lab, uh, is, uh, specifically focused on infrastructure for machine learning for production use of machine learning. So what we're saying there are, uh, view there is that machine learning algorithms are actually already really good for a lot of applications and, you know, new models are coming out all the time. But, uh, what's missing is how to reliably turn them into products. If you just say you have an algorithm that works on a, on a benchmark like ImageNet, um, that's actually a super easy problem because the data is given to you. Uh, the metric is given to you, uh, the metric is given to you, uh, there are lots of algorithms to compare against and you only have to make it work once and then like you can, you know, publish a paper or get a grade in a class or whatever
Uh, but if you're trying to make an application that like runs every night and let's say, you know, decides, you know, what, um, who to give, you know, credit cards to or whatever, uh, there are a lot of things that can go wrong. Um, it's really hard to make that happen automatically. Um, it's really hard to get good data to keep it evolving over time. Uh, it's really hard to monitor it. It's hard to debug it if you, if you see something is going wrong. So those are the kind of problems that we're, uh, focusing on. And we've got a few pretty interesting ideas already. Like for example, one idea I'm excited about that we have, um, some, some papers about it and there's a bigger paper coming out, um, in March, um, is as a means of doing debugging, we take software assertions, which are well known, you know, concept in software
You have these assert statements and we create assertions about, you know, machine learning application. Like let's say, you know, the class of this object you're detecting in yourself driving car shouldn't change over time. You shouldn't think, you know, it shouldn't like rapidly change or objects shouldn't disappear and reappear or stuff like that. And, uh, we showed you can use these to find bugs or issues in ML applications. And you can also use them as a form of training. So you can train a model that avoids failing these assertions. So it's a, it's a different type of, uh, basically, uh, supervision essentially where, uh, you know, you, you can say, okay, I, I have a lot of data. I just want the model not to make these types of mistakes
So that's an example idea. I'm excited about, uh, you know, groups like exploring that idea already. Um, and, uh, it can make a big, uh, you know, if, if it works out and, you know, might not work out, um, it can make a big difference in terms of like how people think about this stuff. I really love it because, you know, that reminds me of just taking software engineering concepts, move them into data science. Yeah. Right. And so what, and what you do with Weld really is exciting for me because I see, you know, folks as Swift for TensorFlow, uh, right. Eugene Brumaka was a compiler leader at Twitter and he, he now works there
And so I think they basically apply compiler techniques similar thing in a way to Weld because you want to output, uh, they have, uh, machine learning to make the presentation project, which, uh, outputs optimize code for GPUs, TPUs and CPUs. Right. And so, uh, I, I think it's really strikes me as the very interesting research that, you know, you can take decades of compiler research and apply them to machine learning effectively now. Yeah. Yeah. Yeah. And there's also actually, I've been doing, if you're interested in compiler work, I've also, uh, been collaborating, um, with, uh, with professor Alex Sakin and, and, uh, one of his students, uh, uh, Jihao Jia, who, um, uh, who's built some really amazing, uh, basically compiler work for machine learning applications that uses, um, uh, things like formal verification or, um, random search to, uh, to, to optimize, uh, neural networks or to optimize the distribution of, of, of, uh, of an application. And, and, and, and those, uh, in machine learning, like in deep neural networks, uh, there are so many ways you can execute them
So many ways you can slice and dice the computation, uh, that I think there'll be a, a, a big sort of boon in, in compilers there. We've got some, um, some really cool results there on basically automatically generating a compiler that's like better than TensorFlow's built-in compiler and stuff like that. Just using the mathematical properties of the, of the DNN operators. Interesting. Yeah. Interesting. Well, this is all very exciting. So we're really looking forward to, you know, great research coming out of, uh, of your lab and, you know, great open source project being advanced
Uh, and we'll be following it. Thank you so much for today. Yeah. Thanks. Thank you.