scale.bythebay.io: Fireside Chat with Ben Hindman and Ian Downes: Mesos, Then and Now
Recording: scale.bythebay.io: Fireside Chat with Ben Hindman and Ian Downes: Mesos, Then and Now
shut about Apache messes this is the first fireside chat we do and in case you're wondering why do you see the picture of fire so at some point I know I was going to a bunch of conferences we should have fireside chats and neither like none of them had a fire and I thought this is a legitimate reason to ask for money back right because you were promised a fireside chat and there's no fire and so I tweeted about it at some point and then I realize you know I have to adhere to my own principle so you know you have to produce an appearance of the fire so this is what we have and I expect a real fire next time and I said well I didn't say real fire right like I didn't say organic fireside chat so so and I know this you know it's been a long day and so we kind of had a happy hour so we we kind of are playing by ear right we're thinking when we do this do some conferences do a talk during dinner right and it feels a bit awkward and so it's an experiment for us so you know bear with us but I think it's really a great occasion right in the spirit of history we have been high man who is the creator of Apache messes and the initial Tim Lee at any endurance who is the current team leader of messes and so I'll tell you my brief kind of history so I was a sort of founder and we did everything with Scala and we basically asked Marius to to mentor us and so Marius was kind of explaining different things the Twitter does and then I heard about messes and it was 2012 and to me it all sounded magical right because essentially he described to me a magical world where machines are not your friends whether you don't know their names that you know there is no develops essentially devups is not given IPS like names and and and you magic and so then you've got this color job and it lives in this magical package with a manifest and you say what it should be doing and it goes out there and it finds like if it needs a service it finds the service and and it connects right and so I just thought this is most amazing why why you know it's been not going around talking about it right because Morris was going around talking about Senegal but Ben was basically I think we you know didn't see been talking about it right and so I think a few years later Ben started talking about it so so I you know I still remember how impressed I was about this idea I think now much more people understand how this works but again right we're at the kind of the cradle of all of this so I just want to to invite Benin and Ian to tell us a little bit more about how this whole thing started and how it evolved internally in Twitter and externally so welcoming and welcoming awesome okay thanks for having us thanks yeah which we started yeah so maybe Ben you can tell us a little bit you know kind of how it evolved how mices evolved a Twitter during your tenure here where did you kind of hand it off to subsequent teams and Ian maybe you can pick up there and tell us how you know where it is now essentially sure so I guess it was early 2010 when we were first talking to a couple people I guess we were really on the meet-up scene when we were first talking about mesos but through through UC Berkeley which is where I was a PhD student we were chatting with a handful of organizations about this thing we were building at Berkeley called Nexus actually wasn't called Mesa as the time was called Nexus for doing cluster management and I still remember the first talk I gave here at Twitter nacked not here actually a couple offices ago I guess did know even yeah not even that building or whichever building was the previous but anyway and I remember I gave a presentation and there was there's like eight people eight engineers this is totally before your time man there was there was eight eight engineers in the room and I I just I remember being super bummed because we've given a similar presentation at Google and there'd been like 80 people to hear us talk and there's like eight people and was like ah this is a waste of time why did we come to Twitter talk and then the chief scientist turned me goes like oh great turnout you got 10% of the company I was like oh yeah that's that's crazy so that was that was early days basically what happened was for any of you that were at the panel earlier mark was talking about there was sort of a lot of early Twitter engineers I shouldn't come early Twitter engineers because early Twitter engineers were like 2008 and then there was her this influx of Twitter engineers in like the late oh nine early early 10 and through who who basically came from Google and so a lot of them saw that that they could get this thing that they'd had at Google called Bork if they could use if they could use Mesa so like oh wow I could use mesas to get Borg and I had had this wonderful cluster manager at Google called org and I miss it desperately and I need it so as a handful of engineers you know some of them by name folks like Bill foreigner John Soros Travis Crawford some of these you know none of them I think are at Twitter anymore although I think bill is still still doing some work with Twitter and so they all said wow could we use this basically to kind of create a Bork and that was early 2010 and then I started to come in and work with the chief scientist at the time guy I need to have your Chowdhury we started to build out was Nexus we soon called it Mesa as well as some other things on top a critical component called Aurora which are rora plus Mesa is what made Borg for Twitter and let's see what else is interesting from a history perspective so I was 2010 the first service I think we ever ran on top was Bill's service started with tea I forget it was not a bird name most services a Twitter bird names and this is like he was like no it was the redirector I think maybe I haven't remembers anyway that was one of the very first services to run on top and then after that we started to get a lot of services to move over 2012 2013 2014 is when ultimately I moved on we just found it a company meso sphere I'm happy to talk about the times in between I don't know maybe you want to pick it up from there and sure so um so my time dates back to 2013 EDD Twitter when I joined if you go back a bit further than that I was actually doing a PhD as Danford and go bears sir go bears so ultimately I dropped out of the PhD but what I dropped that to do was to try and start a company we looked around us when we're doing our PhD my co-founder and I and we saw lots of tons of people who were trying to run the simulations on at that point we had our own our own classes inside Stanford they're trying to run these disolved live simulations most of them were monte carlo simulations so they're pretty easy to paralyze but they couldn't do it they were struggling to even go beyond their one laptop so we founded a company where we tried to make it really easy for them to scout out their applications and run them on something like their own to a private cloud ultimately coming straight of academia we had no clue what we were doing we had no idea about what the market was we tried to sell to academics which don't want to pay for anything and so ultimately we ended up at Twitter a few years later so I joined in 2013 I have to say some of the things that we built in our prototype and haven't even made it enter into production yet very valuable things that for various reasons so I joined in 2013 overlapped with been I guess for about a year or so and then I was an Icee for probably two two and a half years the last two years I've taken over management of this team the mazes team also the Aurora team and some and some other teams are related colonel JVM and things that are sort of coming together but I think it's changed a bit in the last few years so we can talk about that charity yeah yeah I mean one of the things that I thought would be fun I mean so 2013 I can't actually remember when docker was officially announced but so for us 2010 2011 and then 2012 when we were really starting to scale scale containers and do a bunch of interesting things so it was pre docker I select remember were you here when Solomon came okay I'll say all right so yeah I mean so so we were using containers that journey for us actually was we started with lxc does anybody remember Alexi yes we started with lxc it was really interesting because we ended up moving away from Alexi because we decided it was a it was a dying project because there wasn't a lot of contributions to it again this was pre docker we end up using control groups and namespaces ourselves at Twitter it was a project that got kicked off called app app do you member happen so so this is a project called app app that got created an app app I think it stood for the application for bundling your application or building your application to app app and the way it worked is you typed app app I think creates enter and then it dropped you into a shell and then you're like yum install or you know pull down some zips or whatever you wanted to you install some stuff and then you control D and it would pop you out and would give you a zip file and that was a zip file that that the plan was is that people would use that to run on top of top of may sauce because in the early days we didn't have darker and so people were actually they were basically just sending us tar balls or zips or static binaries or any of that and so we create this thing called app app and the security team at Twitter if I remember pay pretty much said like nope we're not going to invest in this there's no way we're gonna let engineers put arbitrary things inside of these zip files and we were like well they can already do that but this makes it even easier and so it got got totally canned and then I remember probably maybe eight eight to twelve months after that and we got this email from these these folks at dot cloud and like hey we're working on this thing called docker we'd love to show it to you since you guys are doing containers at scale and they came over they do this like round table they showed a stalker and I call it super cool but it's never gonna work because security organizations and companies they're never gonna support people to put her inside of your containers so we're not interested and I think to this I think of this day we still know yes so we set up to power from the rest of the community we still don't use darker we still don't use fastest managers for our containers we still force our users to bundle everything with their application and they still run on the hosts file system which is pretty different - I think most of most other users yeah I'd say so we've we've actually investigated at various times moving towards either docker images not necessary doctor itself but the docker image format or the absolute image format and to be honest it hasn't been a priority for us partly because everything we do and so the company is on the JVM so people just deploy jars which deals with most of the dependency issues and in the pain of being coupled to the hosts file system hasn't been enough for us to invest effort and moving away from this idea so you know Mike because you know slightly took a beer so can I ask like a blasphemous question can I can I be a devil's advocate right and I mean can I meet that like I hate containers like why why do we in containers when I heard this first you know good views about Aurora right and the manifest and this color job inside of it right so i another ideal world everything is in scala so why do you need an OS like your new junior JVM right why do you need all this craft around it so you have your your program right and you have some metadata which tells it where to find other programs so there's a service discovery everything is in Scala right and so as an application engineer you don't really need why do you need nos it's basically to me it sounds like this you know the new generation asking what is this floppy disk icon in the email or app or war like I don't know what it does right so so you know you know I did my own share of compiling Linux kernel many many times right and so we have now this atavistic image of an ancient machine it kind of immortalized in a virtual artifact right and we are kind of moving it around so the extra stuff you know the twitter staff are always it's totally fine and it's not doing this right so why can this truth be seen by the outside world why you know people who already have masses why do they feel like they need containers and I start I well I mean so III will say this I think one of the biggest questions I ever got once I first started talking to people about pesos the biggest question I ever got and I got in all sorts of different ways was how do I get my stuff to the Machine and I remember the first time I got that question I was sort of I was shocked I was like I what are you talking about you know we'll send the task to the machine will start the task for everything I know how do I get my stuff to the Machine and I was like what do you mean like well I've got this thing installed at this place and I've got this thing set up this way and I've got this one configuration file that I put at this path and all this stuff and I I think Aleksey maybe the answer to question is because a lot of software is still built in this model where this expectation of the way that we're going to install our software and run our software and configure our software and call it legacy but it's the way a lot of people actually end up doing stuff and to me the container actually just basically makes that a lot easier so if if an organization is willing to go through the I'll call it burden but perhaps it's just an initial burden and eventually it's like a it's constant thing is it constant pain yeah so one of the biggest requests for us is when will we support Daka and the thing that I asked is why do you want daca and they say I want to pull down some image from the web which doesn't go down well with security of course but they're looking for flexibility they're looking to bundle up things are looking to have the latest version and whatever else but when you sort of peel the back this if the organization does make that effort to support something and to guide everyone towards the same platform awesome language River it does give you a low advantages which really mean that you don't need the containerization to the same extent that others do yeah I mean I'd I think if a company just she chooses to go through the arduous task of deciding that they're gonna create the infrastructure so you can just be on the JVM and have everything you need and make that super easy then I think the need for containers goes away but the problem is is just a lot of software out there that is not like that and I think about the number of new projects at each company every day that's trying to take advantage of other other tech they're constantly just want to install it the way that they've installed it for a really long time Thanks the container this is recorder right yeah so I think for us historically we've devote a lot of software internally to Twitter which made sense to run on the JVM I think looking forward we will be adopting a lot more open source in which case we don't necessarily control which languages are written in total so for us on our roadmap is to adopt something like docker images so we can do the sort of thing it just has been do it hasn't been prioritized yet yeah and I did the other aspect that I would I would throw in about that is I still think there's a lot of organizations that are not really like even even if it makes sense to bring in docker containers I still think there's a lot of Hygiene that you could really exercise with containers and that to me is one of the most interesting things is when I go to a lot of organizations who basically are treating containers like giant VMs and they'll you know they'll ask us be frustrated because they'll say well we moved to this this platform so we could launch our apps and you know microseconds that's what you guys promised and then you know you dig down and it's because they're downloading like 2.3 gigabyte container and that's obviously we're not gonna launch that thing a microsecond III don't like containers will go away I know that mark expressed that earlier that will move from containers to server list functionless which I really like that I hadn't heard that one before that was good um I'm not too sure that that will do that at least not super fast but maybe perhaps we'll just get better hygiene within the industry of exactly how we're using containers and yeah I mean I think I mean I understand right on the rational level but to me you know I've been exposed to the higher truth right of Aurora and so that's kind of brings me to another question so when I heard about Aurora I thought like this is exactly what you know in the deal scholar world people need right because if you ask all developer you don't want to interact with develops people you don't want to beg them for for anything including names of servers of things right so essentially the basically this magical world where you just paid jobs on the cluster if they're barely more than the scholar job rights basically you wrap in a thin thin layer and they know where are the things aren't so then Mario described oh this is useful because if you debug your own service like search right you can mark the packets as being kind of your own and then you can dispatch them different alright so so to me that was exactly all that was needed to develop a web scale you know company such as Twitter so so when I've seen mrs. fear take off right I was actually very curious right how this is gonna pan out because what this intellectual treasure encased in in our roar I thought like it's come out right but what we've seen is basically missus fear took up marathon because I think they initially compared to Rails versus spring rights like marathons to is basically with the rails of you know comparator or which is spring and was not open sourced so I want to asking and what is their plan I mean Aurora's still right it contains all of this knowledge how guys are gonna bring this to the world the special since mesosphere is not really using Aurora itself right and bring us to its customers I'm not sure why that's my question and why it's not his question so they're always open so is open sourced it's got a number of active users and companies it's it's it's um there hasn't been the same level of an investment and making it easy to use and easy to sort of take on board it's got a lot of dependencies zookeeper and a bunch of other stuff really it's a question of resources for us we had another time to clean it up and make it useful to others I wish more people saw the benefits or wouldn't see the benefits of Aurora it's it's how did you use it's harder to it makes it makes the difficult problems easier and it makes the simple problems a little bit harder so I think it gives that sort of hurdle but really it's a large scale I think it is what I think is the best scheduler that's out there I'm biased but I mean so I ultimately I think from the Mesa sphere perspective yeah we're we're very happy when organizations have something like if they want to run Aurora we're happy if they want to run run other things as well as I'm sort of mentioned in the panel earlier one of our investments initially in Marathon was because of some of the delay and Aurora open sourcing which was which was unfortunate but it's so creating open source communities is exceptionally difficult and I think I don't think people fully realize I don't know if everyone fully realize that it's a lot of work it's a lot of time and energy that you put into supporting people in the open source and I think my guess is that's probably one of the things that has made it more difficult for Aurora as well as you just need full-time people all the time trying to foster open source communities run events it's a ton of work and I think everyone wants to hack on open source but it's just a tremendous amount of work so even once Aurora was open sourced I think our perspective was yeah you know let's continue to build out a bunch of open source projects and see see where it goes and I think marathon had it had a lot of open-source adopters pretty quickly and a lot more people were using it and so we we continue to invest a lot in there okay yeah make sense I mean it'd still be interest right to see that kind of a path right for for organizations and I guess I think there is this very interesting multi-level complexity right because on the one hand you have open source Apache masses on the other hand you have DCOs which is but a lot of organizations are still doped in pure Apache masses right and then you have scheduler so so I don't know if everybody knows the difference maybe you guys can quickly explain right we have this multiple levels so messes can take either marathon or Aurora right and in this here's something for my workers can explain how these three levels interact so we don't use the service at all I've reload knowledge about that I also have to say that have been for us to be completely candid I think the Aurora male split early on reflected the sort of teams that were inside Twitter I think we had separate teams during this I think measures came along Aurora came along was a super team so it's of architecture somewhat reflects the social structure that we had now though we see mayor's us as being in and really an abstraction for us on the cluster and we don't run multiple schedulers on our clusters we only have one Aurora scheduler on the whole cluster so we don't utilize a lot of the functionality inside mezzos that's around multiple schedulers I think it's one two the key difference with the way that it's used by others to elaborate more on on that sure yeah I mean so the the the architecture of main sauce is to exactly support multiple schedulers so we kind of call a two-level in the in the the maysa world which is to say that there's the first level of scheduling which is meso so it's not even really scheduling it's more resource allocation it's like the it's like the I don't know what the right term is but it's it's it's the really basic how many resources should this thing get how many resources should that thing get research that thing get not given those resources that are available what tasks should run on which of those resources and that was a really really it was a important architectural difference I mean when we first created maysa getting back to the Berkeley days cluster management was not it was on a new topic I mean anybody ever heard of Condor yeah all right so you know Condor is a pretty sophisticated resource manager that was used in a lot of organizations especially financial organizations Condor was pretty simple you you described your declarative specification you gave that declarative specification to Condor then Condor use that to decide how it actually wanted to schedule things so what we were really trying to do with maysa was say that we believe that there needed to be a split between I don't know what you got the business logic of deciding where you want to run stuff and just the platform the foundation managing resource allocation and everything else and like that really was a core part of the innovation if you will in May so sand you know some people they've even asked me like because there's been a lot of other cluster managers that have or resource managers have been created since right there's darker swarm there's kubernetes and there's nomad I think in that orders is the three that got created and people ask me well would you have created meso again if these three things have existed and and the answer would be yes exactly because architectural II we took a very particular perspective of how we wanted to build meso switch was we wanted to split these these two levels wanted to create this this two level system whereas a lot of the other ones that have been built they follow very much the Condor declarative specification pattern and to be very clear I don't think that that's a bad pattern I think there's a lot of organizations that want to take advantage of pattern there's a lot of people that want to take advantage that pattern but not everybody and we wanted to be able to provide a platform where people could do other things and you start to see some of the on some of these other platforms people want to actually take advantage of this two-level pattern so you start to see it in the kubernetes world where they're trying to create what they call operators which I love that name I wish we would have thought of that name for our second level things instead of schedulers because you say they word scheduler and like you scare everybody away they're like oh I don't have a PhD I can't write a scheduler so you know that's like a academic thing I can't work on that whereas operator sounds like way less that's crazy but then you look at how you're trying to like build those things things in and there's like gotchas architectural e to how you'd want to do that on a system where fundamentally you're still in a world of ready to declared a specification and submitting that declarative specification so so yeah I mean that was that was a big part of the way about meso so in the early days of Twitter that was like really new that was like too new like when we talk to a lot of folks at Twitter they're like what that's crazy like it's crazy enough that you want to move us to a world where we're gonna go from running puppet on boxes to putting stuff in zip files to writing in containers and the two level stuff was it was even more new but now it's like it's a critical part of what what we do inside of DCOs it's really really important for everybody that we work with you know fundamentally we're running cassandra's and Kafka's and sparks and elastics and all these things and they are all second-level schedulers and we can run many of them and so it's a really big part of what we do and going forward I think it'll can for the industry I think it'll continue to be a really really critical thing you know comment that was made earlier that I thought was really interesting was Evan said the cloud is kind of like the killer of all these open-source projects and I think the fundamental reason for that is because nobody actually wants to operate this stuff it's like too complicated to operate so they want to use like an open-source key value store but they don't want to have to like set it up and manage it and operate it and even like writing a bunch of declarative specs for this stuff and running it on a kubernetes so running it on Aurora it's still way too complicated there's actual way too much stuff that you actually need to manage yourself so I think we'll actually push farther and farther in direction of these things becoming more and more managed services which is again we know what we're what we're trying to focus on with with the prod and what we're doing there was one more thing that I was gonna say but I totally forgot because I'm a lightweight but yeah I guess I'll leave it at that what was the original question no I'm just kidding so this is great so actually this brings me to you know this question we didn't ask in the architecture spell but I think this is very interesting it's all sounds like masses and other schoolwork of freedom against the faceless giant cloud right because the question you know for smacks tax and other stacks does anybody need them if you have bigquery right like you drop everything to giant Google called and then you essentially you know then it becomes serverless right so but I mean somebody still needs to run all all of this but suddenly the group of people who can run these things becomes very small right and so the question really is right like if we go to this model it will be a sure Google cloud and Amazon and maybe you know $0.10 right and the question is now if this knowledge of managing complex or clothes is concentrated in the very few hands and then you know further further removed from kind of general public how does it connect open source how does it become accessible right so so I'm curious you know obviously you guys should see this challenge right so so what are your thoughts in general who are the customers who would need to run messes right and do you think like in the long evolution what they are kind of estimate how this is gonna pan out against you know giant clouds oh yeah so and I remember what I was gonna say and the last one to it thanks for the thanks for reminding me well so because you just said who are the customers are gonna run meso so most of the organizations that we work with atmosphere it's all about DCOs and the reason why it's all about DCOs which is what I was going to say at the end of the last question is for the same reason that most of you guys probably don't go to kernel.org to get Linux some of you might but it's dubious you go get CentOS or you get Debian or you get something else you get some in prepackaged so in the same way that you could have gone to kernel.org and you could decide it I'm going to use I'm not gonna use Bosch I'm gonna use a different shell or I'm not going to use KDE I'm gonna use I don't even know what else people are using linux desktops right you know you pick something that has a little bit more of an opinion you could swap it out if you want you know you could you could totally grab grab booboo and two and then decide you know what I'm gonna make my my default shell c shell or something else you can do that but like there's some pre back saying so that for us is what is what D cos is and it's meant to just be that consumption the consumption that makes it easier for people to get things so to continue then for your question I mean most of the local organizations that we work with is you so who are the companies that want this stuff a lot of the companies are folks that actually don't want to just be in the cloud or they know they want to be in multiple clouds so if they don't want to just be in the cloud they need something in their own own private data centers to to run that gives them a cloud like experience but maybe without VMs or it's not OpenStack especially when it comes to a bunch of these other services the the sparks and the casandra's and you know the other the other stuff from shrimps back smack and and that's a big focus of our as we keep thinking about other things we want to run so I think recently we had a tensorflow to the mix which is another thing that you can deploy and get a distributed tensorflow pretty easy so I know that's the big focus focus focus I think but I do think what's really interesting is cloud is like one of these really things it's a it's though it's a huge enabler for so many organizations and it's also a huge killer for a lot of open source projects there's a bunch of clouds out there without naming any names that are huge consumers of open source projects but never contribute back to those open source projects and some of those companies then really struggle but you know that the part that's tough about all of it is like we still all benefit because we get message queues by like clicking a button or we get analytics infrastructure by clicking a button and having any get spun up so it's like it's tough right we're all benefiting because we can move faster in these environments but you know it's I don't it's it's not a I don't know it's like a super it's a little gloomy for a lot of these open-source projects so we like to think about it as if we could be a way for open-source projects to shine no matter what cloud you're on on top of DCOs then that would be a good thing right if we can make Kafka run on top of DCOs and then you run DCOs at each of the clouds as well as on prem people get to get Kafka the Kafka project gets to continue to blossom and the company behind the Kafka project can probably continue to push that and and then everybody wins and then the clouds still win too because they are getting paid for they're getting paid because you're getting you're getting VMs run yeah that's it yeah you know I mean actually you know it makes me think right if this company which will not be the name it's not contributing back but used all this open source stuff who's gonna maintain the individual Capone's they should be in the economic model by which they will do some ref share with the yeah we do I was I having a good conversation with somebody about this the other day that there should be a new license so there's like Apache right which is just anybody can take it there's a GPL or if you start using it then obviously you gotta there should be another license which is what what have you contributed back right you want to use the code you could totally use it but how do you contribute back in so many ways you can contribute back you could write a 500 thousand dollar check a year to run amazing conferences for the open source project to support open source developers and that's a one way in which an organization could triple your in some way you're contributing to the open source project and so you get to use the license that make sense so I have a question for you right so I know that you guys did a lot of stuff resource utilization and that kind of brings me back to the history of you know kind of misses fear so I when my sphere appeared right my immediate thought was finally what a fantastic platform for this Mac stack we should know at the time was not this my exact but basically my immediate thought was this is this is gonna be the Arora enable for the masses right so what I described on the scholar world but instead what I heard was any data center managers we could utilize your machines to the hilt so the initial selling was basically kind of pedestrian right like let's utilize the crap out of this right like you pay for the machines you prefer the electricity let's make sure that all this stuff gets utilized so that was the initial push and when I was talking to to flow and Toby I was kind of pushing for the big data agenda and so now it caught on and so I see now this might be a part of the story for my situation super heavy model but so I'm kind of curious so in in mesosphere it seems that this dilemma or kind of duality resource management versus pipeline enablement right it went finally the kind of two pipelines and but it sounds like a Twitter you're mostly focused on the resource utilization can you talk about these two heads of misses yeah so the workload that we run is sort of mostly just all the other stateless services so state being somewhere else not that they don't have any state we don't run a lot of I don't even sure the smacks deck is can you fill me in this is spark my Osaka Sandra Kafka okay yeah so we don't run much of that and that's all up on somewhere else on some other classes we would like to run those on the buses I think what you're asking about is the resource utilization are very existing clusters so yeah so for us mez let's have been has been pretty instrumental as escort NC has gone up we've got services inside the data center there to run on a single machine and so that's okay if you're running on a quad core machine but now you can't really buy a quad core machine now we're buying sixty four three eight machines whatever else so being able to take that service containerize it put it on and run it somewhere in our platform taking up only the necessary resources that it requires that is what is pretty instrumental to us and achieving high reservation so I'm hesitant to court it actually you do I solve utilization because we do attribute all of the threads on the box to a particular container but that doesn't necessarily mean that we have high utilization of the box itself that's a separate question I think I can talk about that yeah so for us for stake the services most of ours are actually latency sensitive so we're fairly conservative and how aggressively we pack stuff onto our onto our machines but it's really an area for us where we have you know classes with many many cause in the cluster how do we increase the utilization overall of that cluster because even you know even 5% increase tenderization translates into a lot of dollars for us so there it's pretty important and we've you know we found that and this is through no fault of this is not at the Maysles level this is at the individual scheduler level hours of experiences at the Linux kernel in its stock configuration is not great for you know high utilization workloads that are that are also latency sensitive so a big area of work for us is how to improve the resource isolation Tiger being good predictability low variance in performance for the containers but also being able to trade off between different workloads running on the same host and our overall increase in utilization of the cluster I don't know how much of a concern that is for other customers who are operating in at smaller scales but for us it's it's one of our leading concerns comfortable yeah I wouldn't even say smaller scales I would just say earlier in the process you know I'd us about patterns in the last last panel and a lot of organizations that are first doing this transformation in their companies I mean it's just so early it's so much earlier whereas Twitter I think is at a place where they can actually be doing a lot of interesting interesting stuff because they've just got the platform there so do this atmosphere do you now kind of look at Twitter as the carrot of innovation right is I mean is there just like flow going back from Twitter to mesosphere for certain things I mean if you look at if you look at the talks at mesas con that Twitter giving there are always things that other customers of ours or is like oh well we want to get that thing as well but but it we're sort of like well you're kind of a ways from there you get your clusters to the person some of the other stuff yeah thanks and you know one thing I mean I learned recently and I can say this because I mean I'm not possessive eater you know that I don't like apple and over are running huge mrs. Kloster so and the questions who has the biggest mrs. Kloster of the world and the question was very interesting to me that it's not even an open like it's not a salt question right like it's it's still you know we don't know if Apple runs a bigger class or Twitter or I mean at least I don't all this and so so to me it's very encouraging right I think it's a very good sign that these kinds of companies selected messes and and I did know actually this until I knew that Apple is a huge user of Scala and substantial participant participant in this conference right so I kind of learned that there they run one of the biggest message clusters after they enforce right so to bees are encouraging so maybe you know like it's getting bit late so maybe look I'm wrap up with their first chat with the question where do you guys see messes coming going inside Twitter and outside like what their plans are and how they're gonna collaborate maybe going forward so right now we're running this takes all the sexes of workloads what our sort of gold for our team is is to be the single infrastructure layer across all of our other hardware so running all of our services on top of of mezzos not just eight lists but also all about stateful workloads all the backing stores even going as far as Hadoop and everything else we want to sort of basically run everything on top of us and that is a list to do with resource utilization for other services more to do with I think to your earlier point people were really sick and tired of managing hardware and they want to get out of their business completely they want to think about this is my service these other resources that I need go right up for me and just keep it running so we want to offer that for the entire company for all services that we need to support and we work closely with so I've made a few folks to to be able to do that they're really pushing the envelope in terms of stateful frameworks we'd like to adopt some of those ideas into Aurora and take advantage of the work that's been done and mezzos that actually enables all of that yeah and I mean I mean at as you were talking about scale at Twitter scale and some of the other companies you were talking about it that their scale we learn a tremendous amount about what we needed to do to improve the project and improve the product we have a I think if you just even if you just take those three companies you mentioned and you maybe you throw one more in there which is Netflix there's no that's a lot of containers so that is running on top of may sauce and yeah we learn a lot from from people doing that I mean I know if there's any mace those users in the audience all right sweet great they well I said when you gets a certain scale like there's just certain things you know not to do like hit state dot like hit slash state which is an end point we have some uses some of our customers at the company decide they want to hit state or JSON and it's not pretty it's really bad like at the scale of some of these customers just like it takes like minutes to actually do that and it like completely blocks the master it's like crazy so it's we're working on a fix for that though just yeah really good we've just one other solve an issue a while back we reached a certain scale point where on a master failover when all of the of the agents who were listening to zookeeper when they tried to reconnect the master actually determined that it was a sudden flood attack because there are so many incoming TCP connections at about the same time at which point most of those were denied and dropped in everything else and of course they were try to reconnect again and it would just keep flooding flooding and eventually take down the master it would fail over and this would continue that's a big cluster yeah so but we contribute some fixes back and things find now yeah so I that's a big part where we where we still worry so partnering a lot of a lot of open source folks and it helps make the make the tech even better and then people who end up adopting and as they grow themselves as they scale themselves they get to benefit from from that work being done because we've got some really big large-scale users so that's how I'll keep collaborating with folks like Ian I think we have lunch scheduled soon right within two weeks you said yeah all right guys this I think this was great you know one thing I learned personally you know in my mind if I replace scheduler with operator I'm like I feel so much cheaper now so I'm gonna do this from now on you know it's never it's better late than never right let's just all start doing this right we'll all feel better well it will become clearer so its operator so thanks a lot this been a long day you guys are the veterans and you know so this is awesome and [Music] [Applause]