scale.bythebay.io: Fireside Chat with Evan Weaver and Boaz Avital: Storage, Then and Now
Recording: scale.bythebay.io: Fireside Chat with Evan Weaver and Boaz Avital: Storage, Then and Now
so welcome everybody to the last event of the day it's been a long day I think it was a really great day if you've been here for the panel who enjoyed the panel oh yeah like we do not happen like this if that's you know that was Vitale's idea and one of the reasons we you know he proposed this idea that I added too many people to the panel so he wanted to have about four people on the panel but when having six people on the panel he said the only way we can do this is by debating so we have this amazing debate right like if our program is good for machine learning so so that's really awesome so that's an innovation in the fireside chat the innovation yesterday we had such as Benjamin Ian Downes and I think we really learned a lot about messes and mesosphere and and you know why we're all hit containers so that was really interesting so today we're talking about storage and so we have similar format we have Evan Weaver who was the first lead on storage team right leaders as strong word for what was happening at the time and bas is the current leader of the storage team so I think you know I love you guys introduce yourselves and tell us basically what you were doing them and what they're doing now I may be a boss for you like what you were doing at the beginning when they joined Twitter what you're doing now right and and how it all relates to storage yeah I'm Evan Weaver I was Employee 15 and Twitter and started in 2008 stayed about four years and initially it's just me and I'm guy named Rayleigh pointer work you know and scaling my Segal and that grew and grew into what we called the infrastructure team and eventually there were multiple teams including the storage team which Boaz it's the cultural inheritor of and we built all the distributed operational storage so that was tweets timelines users storage social graph image storage the cash some other storage those were effectively point solutions we were frustrated that we never had business like investment horizon to create a truly reusable internal data platform for Twitter or for the open source for everyone so when we left a number of the people who are on my team eventually came with us we started a company called fauna which is making fun of DB which effectively is that platform but from a vendor perspective and over youto about what happened after that within Twitter my name is Boaz avital I joined Twitter I want to say two months three months after you left and under the storage team at the time we were working night and day making sure that the databases stayed up the user databases that you guys built and the the Cassandra database is serving data different kinds of data that we created that hawk I mean over time we the company continued to grow the number of users continue to grow the number of things that we were trying to build a Twitter continue to grow and multiply and the the team that handles storage can't grow proportionally with the number of things that the company is doing so eventually we decided to build our own database with a lot of things in mind that that were important to Twitter at the time given our constraints and that's things like multi-tenant database that has self-service as a first-class consideration operations as a first-class consideration so that this thing can scale truly scale in practice not just in theory to the scale the Twitter needs both in terms of requests and storage size but also in terms of the number of people using it and the number people operating it and that's databases called Manhattan today it's we started building it like a fume six months less than a year after I joined six years ago and and today it's the canonical store for tweets for DMS for users advertising data and a whole long tale of things that I don't even know about because it's self-service and nobody has to ask so I mean the immediate question right as an outsider so we're in the MIPS talking about open source about people picking on open source projects and using them right and so one of the examples max tech and sees Cassandra so why is it not good enough for you guys right so like why couldn't you pick one and you know so if you look at the edge base right so Facebook picked HBase made huge contributions but it's kind of all out there so I don't know how much it so they did he'd rather they did like some internal additions right so but so so you seem to follow kind of the opposite models so for them Gilda's all database which you're gonna manage and essentially the same internal internal manage database so I'd like to kind of understand better so are they kind of descendants of Cassandra right so you talked about early days of Cassandra and and how they different from Cassandra and how they're different from each other and why do you have to kind of frame them the database for their specific needs you have a twitter or you for the customers I mean we were involved in the early Cassandra community and in a way you could conceptualize fauna as like a Cassandra 2.0 but I think I think that's doing it a disservice because a lot of the things that we wanted to standard to be at the time turned out in practice to not be the right set of trade-offs especially in transactional operational data so a lot of the things that Boaz mentioned as being important to Twitter later at a later stage of mnsure the same capability that we also discovered both in our experience at Twitter and then also in the marketplace this enterprise software vendor like multi-tenancy QoS management a better security model strong consistency global replication and data sovereignty control a whole host of you know especially for a start-up as chaotic as Twitter in the early days things we couldn't even dream about solving before we saw to be immediate scale problem and then keeping the site running problem so I think in a way it's that experience in my view which is translated into both systems not the code itself but you're welcome to sure I mean we we were really heavily into Cassandra it wasn't the storage system that was storing kind of what you would think of as the main Twitter data the social graph and and the tweets and all that but it was a great system and it was the easiest way for us to store a lot of new use cases and a lot and to scale a lot of use cases that at the time there were no other systems Cassandra installations that scaled to the level that we did including like millions of of Rights second for our observability stack but the change is ultimately it's a very hard decision and we had Cassandra committers and on the people on the Cassandra PMC and we had to make the decision about whether we wanted to to contribute back into for or to fork or to write something from scratch and and we're working on in this company under the kind of the reality of of all the constraints that we have to work under in order to deliver value to Twitter and make things right as fast as possible and the number of things that we identified at the time as issues with Cassandra that we wanted to change to simplify we had different goals in terms of scale versus features we different goals in terms of portability we have this wonderful infrastructure that we can leverage as being part of a large company so there were a whole host of things that we wanted to replace in terms of internals that that didn't match the road map of the community at the time and we didn't want to be disruptive to the community either so so ultimately that caused us to to decide to build our own thing because we knew what we wanted and and had to do that quickly so I want to kind of still kind of stable this topic for a moment right so we talked about this Mac stack it's kind of one of the no terms generic terms right so if you look at the CC it doesn't mean cassandra mean couch by is basically some persistence it looks like it's the weakest late letter so keys the strongest cuz Kafka stays right like we're gonna talk about architecture Kafka stays there is everybody's using Kafka it's very hard to find a company you know which needs a message bus which is not using Kafka right so animals only have kinases right but basically it's two key so so somehow you have very solid key and S stands for SPARC it's really hard to compete as part you need DTLA you spar right messes fees for operations again we have you know uber netflix apple peak masses you know without missus fear right so Ark is an API again this open source what is it about the databases which makes them first of all so hard right and so organizations and so you mention some of this right but I'm curious is this like a law of the land that big companies have so spoke needs right they're not reinventing spark they me that's called English they kind of you know center of the world but you know comfort is not reinventing message buses all the time so LinkedIn invented one you know and now people are just picking it up so so is it something specific because date is at the heart of the company it date is very different what makes this you know piece so important that people remain their role I mean we did write several message buses at Twitter to be fair and I think anybody's a bill dilemma mister I think I think I'm sure katka is used now along with whatever kestrels like grandchildren are but I think ultimately data especially synchronous transactional data is simply the hardest problem to solve and that means it takes the longest but at the same time there's the most opportunity to make an improvement on the status quo like I mean when we started at Twitter it was my sequel vertically scaled and not very not very high in terms of vertical reach there were many iterations since then that succeeded and many that failed to get to the point where you have a system like Manhattan um and I'm sure I'm sure you can tell there's still a long way to go in terms of capability in business value that can be delivered there and I think fundamentally it's just easier not to try in a way and stick with the devil you know either for sequel or now people now some of these no Seagal stores like Cassandra are incumbents but their fitness for purpose continues to get worse and they continue to fail to keep up with the way people build modern applications and I think if you can actually build a genuinely modernized system the way Conklin kept that did for for message buses there's huge huge value there so that motivates people to keep that in I mean especially early early on years ago if you look at what a database is like it's not just one thing that anyone was trying to get out of it obviously everyone would love like a my sequel that scales forever and you never have to think about it then we wouldn't we wouldn't be having this discussion but like the trade-off space of what are all the things that you can build and if you choose that one thing is important to you you have to give up something else is huge and there's points on every point at and any database you choose is going to make different trade-offs and no two databases like think about the data model the same way or how requests happen the same way or how replication happens quite the same way and those all have consequences so eventually you get into a position where what is the main way that your company operates and how do you leverage the database to be the most useful for that and and when you guess it was earlier web days when there were only a few companies if there was one of them where we were at this kind of enormous scale you run into problems that people just didn't run into thank you so uh again I see some similarities right so fun is the manage database for end users for multiple companies and madness essentially managed database for internal suitor so can you talk a little bit about the managing aspect how how does this happen so you know I'm a new company I want to work on I'm a new team at Twitter and describe like what needs to happen for this so fine your phone is a monolith you run one jar on each machine in each data center in your cluster VM instance container or whatever you have but the user experience even whether you're running on premises or in our cloud is essentially a SAS experience you Coria API to create a new database there's no preventing capacity it gets a priority assigned potentially by you or potentially by an operator on the other side of the the wall so to speak that limits the damage you can inflict on the machine resources with your workload but it's an like i/o cake you know data fabric or utility essentially um that's what that's what we deeply and painfully missed when we were at Twitter and that's what Manhattan in a lot of ways have solved I think the biggest difference between the two systems because fauna was built to be packaged and shipped to an operator and Enterprise who we as the vendor don't control it's all in one whereas Manhattan is in my understanding designed around the service architecture that Twitter uses for everything and you can talk more about maybe the trade-offs involved in doing that sure and yet I mean when you when you build something into the infrastructure like that you make a lot of choices that are simple for you but but make it hard to make it portable later we want to open source it up like that and that's unfortunate but but it doesn't make a it gives you a dependency on other people on other teams as well and other things that the company have to be running for the database to run which when you're at the bottom of the stack like makes you a little bit uncomfortable sometimes but ultimately it's a it's a huge boon to the speed of development and and the things that you can rely on for us to answer your question it's the kind of the the the thinking was if you were developing software outside of Twitter you'd be putting it a native OS or in a cloud and you'd be able to use a managed service ID to be service and just ask for capacity and get it and it seemed silly that it would it was harder to set up storage inside of Twitter which is this walled garden where we all trust each other and like it's a there's a lot less controls and then it would be getting it outside of it and so we built it specifically for that need to lower the amount of time that it takes for someone to get storage going so that they can think about how to actually build their their services instead of thinking about provisioning databases that was the goal and like there's a lot of things around it that we had to build and metadata distribution and api's and webview eyes and all this stuff that you don't normally work out if you're working on a database team at a company but getting that breadth like thinking about this database as a product truly like we're a start-up inside of Twitter instead of just just building a technology and in managing that product and running it ourselves really changes the way that you think about building the software and you build it in a better and more robust safer way at the end of the day yeah I agree with that like one of the things there was a big difference for me going to photo from Twitter it was sweet at a marketing budget so you know we could spend time telling people how and why to use our software internally at Twitter when there's always a next fire to fight and we here's some code you know if you have to scale data like our code is there or you can put it in the existing system but and don't break it cuz there's no isolation so don't put too much or you know there's this special pool to plan or all the bad actors and like you say very little leverage especially for the product organization and ultimately extremely efficient but also extremely brittle systems and both both platforms are a response to that one thing that I envy though is that when we get customers we're not allowed to charge them and turn a profit and so and so they're not thinking about it that way we have to we have to break even on our costs unfortunately so yeah can you talk a little bit about running you know operational component right so do you run on set on messes we don't run my hand no missus we're running on on on physical machines okay so so basically so both systems have to assume that the cluster is just a set of machines like you do your own Klosterman your own resource management yeah I mean traditionally even in non distributed databases you treat the operating system is no more than a collection of device drivers like you want to own everything end-to-end because you like tight coupling is mandatory for performance and phone is built completely on the JVM it's implemented in Scala it runs on a standard JVM which makes it easy to deploy and you know what else happens outside of that world is not specifically our concern our goal is ubiquitous and then relying on the JVM to deliver that maximize control over the machine resources that are there because a lot of a lot of this stuff you have to do and then we had to do in these systems and Twitter for performance reasons required breaking down interfaces between systems whether that was libraries or services or no processes on the same machine if you can't share information you can't make better performance decision so the more that your own for a database the faster and also more predictable performance you can have which is critical for customer facing workloads and that story was the same for you know a system R and it's the same for us today so how do you think about performance and funnel so our focus is on predictable performance first and then throughput second so it's it's it's it's paramount to be predictable after that you want to be reasonably low latency and after that you want to be as high throughput as you can within those constraints it's worst you have a system with twice the throughput but twice the variance because the applications get built to the imply and performance profile that the system delivers not to whatever like your Doc's say it may do in a worst case scenario so it's really about you know if you want the the customer whether that's internal or out to have a great experience you need to make sure you don't leave them astray in either direction in terms of the performance profile and it also means performance under faults which is like the traditional downfall many distributed data systems where it works great in a lab or in a vacuum but you know under real-world varying workloads and with faulty hardware especially hardware in the cloud which is even worse than premise hardware or even to even to this day like it's great you were up all day but that one minute we're latency went up a hundredfold destroyed any perceived trust any of your customers have in the system I mean we certainly we there was an explicit design goal in Manhattan from the beginning optimizing for worst-case not average case performance and and and over time I think as Twitter's architecture matured we what what used to be like a short blip in latency for other systems causing a lot of problems up and down the stack stopped being the case which is which was a great benefit and one of the reasons why the fail well kind of went away and so I totally like agree with a lot of what you're thinking but also something that the things that we continuously running to and that I can't imagine running having somebody else running my database on their machine is constantly we're having like kernel level interaction issues other machine level configuration issues things are things aren't set up correctly on their box and and like and it's causing pathological performance for whatever reason and we have and we've had to investigate that ourselves and become you know with the help of all the great like OS and VM teams our Twitter become experts in how to run these systems in reality and just handing it to someone else I'm I can't imagine like oh you guys are running this bad version of this kernel with this bad version that's OS with this thing that we try to do and so yeah your performance is crushed I'm sorry like have you run into anything like that yeah we have and like that's why you have a relationship with your vendor if you want you know beyond adequate performance under the worst-case operational topology and you're running it on premises you have to someone has to investigate and to end understand that there aren't artificial bottlenecks induced bike figuration or dependencies because you know there's still like chips and an operating system and other stuff going on we can we're not shipping you an embedded appliance at the same time though that's one of the biggest benefits of us running a first party cloud and - and deploying to that cloud and like our cloud is layer it's software-defined it's layered over AWS as well as GCP and one of these days as you're as well um we experienced in the diversity of deployment environments we can do that management and an offer to you the cloud customer consuming fauna over an API arm that perfectly tuned and safer deployment scenario but if you're the enterprise and your compliance requirements there's no way we can possibly address all those in a minute service so thank you so it's I mean it sounds like there is a lot of similarities right like it's kind of another similarity so I noticed on the Twitter panel you were the only one arguing for the closed-source solution right and I think right so so and and Manhattan is not yet open source right we have this meetup here maybe a couple years ago right then I was me - asking you know when is this time gonna be open source and I think that the answer is very very cautious right and so you said like you guys were closed for six years so do you plan to open source no heaven at any point and with what what we're thinking I say about open sourcing such a complex thing as a database sure I mean if the it's not off the table it I would love to do the community around a hat and I see value there when you're thinking about this complex system that's so tightly tied into our infrastructure and where so much of the value comes from the ecosystem that we've built around it the good tooling and the good agents that manage things and the and the api's that run alongside it and stuff like that that it just has has to be done mindfully and that's work like any other work and we're here working to make Twitter better primarily and so and so that's work that has to be prioritized against that yeah I mean our experience working on some of the very early Twitter open sources the calculus dramatically changed even before you know we were six hundred people even when we were two hundred people we're in particular for the thorniest problems which were often the data scalability problems the help you got from the community started to fall below the effort it took to maintain that community and part of it is because of this need for specialization in your own system versus trying to address more general concerns for people who are solving much in particular problems and much different levels of scale and I even experienced this and like much less significant open-source projects like the memcache driver I worked on a lot where it was in the critical path of Twitter's rails app for years and years and any like a single-digit percentage degradation in performance was a crisis but I get all these patches which would be like let's add a feature makes performance 10% worse that's not a big deal it's Ruby I'm like it's a big deal and it's just hard to balance all those competing concerns and at the same time I think in particular like things like containers and other middleware which derive a lot of their value from conforming to standard interfaces in practice for databases that's never really been the case like even sequel isn't truly a replaceable standard like there are edge cases and functionality there are all kinds of edge cases and performance the operational interface is completely uncontrolled like there's no world in which dropping something new in gets your benefit and that makes it very hard in particular to get a lot of people to collaborate each of whom are narrowly focused on solving only their most acute problem you work with a lot more other companies than I do obviously but my perception has been that definitely early on very few start and small companies had had what we would consider like real scale but but it seems like that's changing and like a lot more small companies are finding sources data wherever they're coming from in the in the millions of requests per second range with varying levels of latency requirements do you also see like more companies are using that big scale yes but at the same time you know Moore's law has for better or worse continued to keep up and the the level of vertical scale you can get in the cloud on demand from vertically scaled legacy solutions has essentially kept up with the pace of data growth so I don't know that the equation has changed that much because you still have very few companies which are forced to the way Twitter was to invest in scalability for its own sake well you do have our companies that suffer because of the weight of integrating all the different data solutions they're trying to use get captive by overly narrow either no sequel or legacy scaled Seigle solutions and then get to a point where their product their scale is okay but their productivity has ground to a halt because they can't both expand and isolate the different workloads they're trying to run against the same data and we had that problem in Twitter and it's still happening today so I think like scale itself arguably never was the most acute general issue but productivity at scale absolutely was and continues to be so thank you so I think we're coming to the end of the half hour issue held right so I kind of want to to end with the question what is the hardest challenge for you right now rightful forma and for Manhattan why what keeps you up at night and what do you think will happen when you will overcome this challenge I think we have a really interesting challenge a Twitter building building our database because because as a company we kind of have grown to rely on it there have been times when I've I've actually told someone we offer my sequel and you know for people to use for their data I look like anywhere else shredding and I've told someone like your data model fits better my sequel your requirements fit better my sequel you should probably use that and they said I'm gonna say yeah but Manhattan is so easy to use and I'm already used to it and this metal fit the no sequel model so people try to use this database for all kinds of different things and and like Evan said the more you kind of understand the use case the more you can fit your database to it and so in addition to making all of our the core things that we do better we have a really interesting challenge to work with all the different teams that use Manhattan to Twitter and and build new features and a new ways of storing and accessing and querying data so that it continues to grow and be useful for everyone regardless of the kind of ways they need it in a way I have the same answer I mean technical friends can I went first technical problems can be solved a changing behavior is the hardest thing and in particular in databases and transactional operational databases like the risk of using something new whether that's something that's not Manhattan and Twitter or whether it's fauna compared to whatever the incumbent thing maybe is rightfully perceived to be extreme and both showing people the value they can get from working with a new paradigm and a more trustworthy piece of infrastructure just takes a lot of work and it's people work it's marketing work its sales work it's documentation and support work as well as actually delivering on the core premise of the technology itself but I think that's always been a bigger challenge for anyone in the day of space then making like a algorithmic improvement or something like that Mona Mona doesn't go down so I'm not being woken up by pages or anything so at Twitter that was my main thing that literally kept me up at night but that's no longer the problem for us it's a lot better great well thanks guys I think we wish you know the most of successfulness as a company in new company in the market place and now Congrats so to fundraise and be presence here we always support scholar companies in other companies and boss you know most of luck to you at you know storage leader I would hope at some point you know see more button happen so very interesting topic right so we're happy to have a meet-up on this and looking forward to learn more thank you very much thanks guys let's give a round of applause or thank you very much it's been a long day we have about 15 minutes and I don't know if anything is left there but you know enjoy yourselves and I see you guys bright and early tomorrow