Devreal

SF Scala: Evolving Infrastructure After Twitter: a Panel

SF Scala: Evolving Infrastructure After Twitter: a Panel

Recording: SF Scala: Evolving Infrastructure After Twitter: a Panel

. Folks doing Twitter open source, I think it's super exciting. I've met Evan actually a while ago. We talked about Fauna, kind of came on my radar through Twitter somehow. Everything about Twitter comes through tweets. And so I was really intrigued by this and I met Evan in his previous office, which was this house next to Stillman Street. It was super exciting, so I wanted to find out. And Mark actually was on our meetup several years ago and then completely disappeared

And I was kind of seeing him on Twitter, so I'm super happy to have him here finally being able to tell us. And William is a big friend of the meetup, so he came to Saskala many times in Reactive Systems. So it's kind of very interesting. My overarching goal was with Twitter people to actually bring their awesome work into the world. And because we have Arca, which is comparable to Finagle, but talked about much more. And there is Twitter developer tools which are awesome, but we are not seeing them outside. There is databases. So I think one of the topics I'd like to kind of cover is what can we do to kind of bring all this goodness out

So I just want to mention this. I think we need to do more work. We did Scala, by the way, inside Twitter. So I thought, basically, I will resort to drastic measures. We'll bring Scala inside Twitter. We'll absorb everything as a sponge and we'll take it out. So I think we succeeded a little bit, but I'd like to understand, you know, what are the use cases for Twitter's open source? What are the existing customers? What else can we do, basically, to evangelize? Now that you all have different companies who have to do it. So the way we're going to run this, we're going to let the panelists introduce themselves and their companies and their stacks

And kind of make the key points, what is special about their product, why you should use them, right? Or why the customer should use them. And we'll kind of try to see the connection, right? And kind of, we'll try to see, like, what is the theme here? What is Twitter infrastructure kind of learnings which people on site Twitter now can use through this new offerings? And we'll kind of run with this. Folks can basically have conversations inside the panel. I'm not going to ask you, you know, every question in a row. We can do that. But, you know, I hope you guys can strike conversations and follow up on each other's themes. And then we'll open it about halfway to the audience questions. Does that work? Cool

So then let's start with introductions from left to right. Hey, I'm Evan Weaver, CEO, founder of Fauna. We're building an adaptive... Thank you, William. We're building an... We pay this guy. We're building an adaptive operational database. Our team is from the early Twitter infrastructure team

I joined Twitter in 2008 as employee number 15. And we ended up building out the distributed operational storage specifically for tweets, timelines, the social graph, image storage, the cache. Yael joined our team a little while later. And some other storage that I don't remember. And some of those systems are still used today. At the time, we weren't database developers. Like, I was a Perl hacker, basically. We just wanted to scale the site

So we had to learn everything the hard way. And leaving Twitter four years later, me and some of my colleagues from my team there were surprised to find that this problem of, like, having a database which can scale but is also flexible for application development remained unsolved. You could get, like, a key value store and it could scale but it couldn't really do anything. Or you could get a document database that didn't scale. And we figured if we don't solve it, nobody will. So we started wanting to build that system that doesn't need to make trade-offs on the path from growing your product or your company from small to large. Hi, I'm William Morgan. I'm the CEO of a company called Buoyant, which is also comprised of a lot of different things

Because of a lot of ex-Twitter folks, we build open source infrastructure tools and we focus on people who are moving into what we're calling the cloud-native stack, which is containerized, built on orchestrators like Kubernetes and Mesos. And we, of course, you know, we saw kind of this migration that happened at Twitter, a very public migration from, you know, this monolithic Ruby on Rails app to this big, massive, multi-service, highly distributed application. A lot of the problems that came with that, specifically around service-to-service communication. So you heard Alexi mention Finagle. That's something that we've tried to take. And actually we're focusing, and this is maybe interesting from the Scala perspective, we're focusing on purely the operational side of things and not so much on the programming model or on the developer side of things. And I was, I guess, not quite as early as Evan at Twitter, but early enough to see a lot of pain. And, uh..

How much of it did I cost? I don't know. I was happily several layers above on the stack. So, you know, I assumed there was a database somewhere in there, but, you know, I didn't know. No, you knew. I didn't really know what happened once the bytes left the service. Yeah. So, uh, I'm still at Twitter. And also I'm not a CEO of any company, just so that we're clear

Yet. Not even half time? No. So, the kind of, uh, so I have been, not only am I still at Twitter, I pretty much have been working on the same problem domain since I joined in 2010. So I think I was employee number 300 or something, something in that range. But if you remember, that was the time when fail well was a meme, right? And, uh, and my particular role in that is every time I did anything to any cash, um, there was a bunch of fail well, wells getting thrown. So that was the state of Twitter infrastructure at the time. And one of the reasons, uh, that I've been on the same project or same team for so long is, uh, cash continues to bear the brunt of the consequences and the prerequisites of scalability. So, in many of the very, uh, essential Twitter services, we continue to serve like 98, 99, 99.5% of the traffic out of cash

So if anything happens to cash, basically you can guarantee that the database crumbles under the weight. That it doesn't have to lift usually, right? So, so this, uh, crucial role to scalability and also just the, the scale of Twitter infrastructure means we get to see a lot of the traffic. Uh, you see a lot of the dark corners of, uh, high performance distributed service and data centers. So there's a lot of wrinkles you can iron out if you look at it hard enough. So my gig, uh, Pelican cash basically is what can I do first to fix all the problems that's in the existing public offerings, Memcache, D-Redis, that people generally use. And also if I want to be productive, uh, I want new features in cash. How do I get that? Do I go talk to the Redis developer? Do I go talk to the Memcache team maintainer? I don't want to do that. And so I want to give people the flexibility of creating something that's scalable and also flexible

And that's the project that I've been working on. I'm Mark McBride. Uh, I joined Twitter significantly after that. Yeah, I was employee 80. So it was like a year, probably. Uh, I started working on the streaming APIs myself and one other developer, which was, I believe, Twitter's first large scale Scala thing in production. Um, I don't think anything preceded it. Did, did Cassowary come first? That was after

okay. Then I think it was. So the, the, the streaming API was, I think, at the time, you know, very dense service by Twitter standards, handling like tens of thousands of connections per box in a very sort of old style thread for connection model, which was interesting to scale. Um, I moved from that to doing a lot of the front end service migration from rails onto the JVM. And so where Evan's team did a lot of the, the back end migration of splitting things into services still behind rails as a router in front. Uh, the splits of taking actual customer traffic off of the rail stack onto the JVM came later. So I worked on that. Um, and then I moved from that to working on build tools for Scala and then I quit

Um, that job turned out to be real hard. Uh, and so. I thought you found your true calling. Oh, no. So, so I, I went from Twitter to Nest where SBT was still in place at Nest and I was too tired to actually even argue with that. Um, and did not. And so I did a lot of, um, Nest server infrastructure was all Scala and I think still is. And so I did a lot of Scala networking and sort of device communications instead of customers, which is very interesting too

Um, and really at, at Twitter a lot of the tools we built to help with that migration of customer traffic from rails onto Scala, um, required a lot of new techniques that we built out along the way with fits and starts. Um, and, at Nest there were a lot of times we wanted to evolve infrastructure and just couldn't because there were all these devices out there that were fairly fragile. And any changes caused, you know, a very high likelihood you'd have outages on devices that people couldn't get too easily. So we didn't, um, and the infrastructure stayed static and that sucked and there was a lot of technical debt. Um, so leaving Nest, it was really, you know, similar to, to Evan, um, and, and William too. I think looking around at what problems did Twitter solve out of necessity that you looked around in the market and aren't solved because people don't have time to just build them as products. Um, where could we fit in? So we're really looking at, um, sort of a stack agnostic way to help people manage that tension between moving fast and keeping customers stable software releases and migrations. Um, maybe, you know, uh, we don't have to do all of these questions, but I think the first question is really good

Which Twitter word story exemplifies a company or project? So maybe you can kind of tell a word story which showcases why, you know, your project exists and what problems kind of it's supposed to solve. In any order. I think we should start with Evan. All right. So, um, when I joined Twitter, the first thing I did was, uh, add a timeout to Apache. Um, because before that, if the database, which was a single, my single database was too slow, your browser would timeout after five minutes and you'd see a white page. Um, well, if you add a timeout, now you have to have something to show. Um, so I was browsing around in the public images folder

And there was a whale there. And, um, Britt and I, and Britt is here actually, uh, were like, does this whale make sense for this page? And we're like, sure, why not? Um, I don't know how it got there. Uh, that was a great mystery. I think Jason Goldman had bought the stock photos for some, some reason, lost in time. Um, but we added the photo and rolled it out. And Jason was like, I think the robot is better, but it was too late. Uh, and that was the beginning of the whale. And I spent the entire rest of my career there fighting that whale, um, including building out the team and building all the systems

And like the fundamental issue there is, uh, a very complex interaction between scale, overall efficiency, and the rate of change in the system itself. Like, at the time, Twitter was one of the first companies to, um, suffer and benefit from mobile-first adoption curves. So we had people coming on really fast. It wasn't like Facebook where you could partition the network and go college by college. Um, it was post-broadband penetration, uh, the beginning of smartphones and mobile adoption and heavy SMS usage. But it was also pre-cloud. So, like, if you wanted to get machines, you had to call the colo and, like, their salespeople would have staked dinners with HP's salespeople. Like, six months later, you'd get some hardware

It was immediately redlined. So we had to both improve efficiency the whole time while also solving these scalability problems. And ultimately, we did that over and over for all the low-level, um, canonical business objects. But we couldn't get ahead of the curve and build something that was a reusable platform. We could build a partitioning scheme for the social graph, and it had some tweaks. And then we could basically copy and paste that code into the new system for the timelines, for example, which had a different, more cache-oriented backend. Um, change it, adapt it to that workload, and fit it to what Twitter's data problems were at the time. What we were never able to do, and which Twitter is still working on now, is build the platform that could service all those workloads in a coherent and composable way, single piece of infrastructure, um, let other teams iterate on what the product was without having to incur the burden of those scalability concerns

Because, like, a new project would come down the pipe from the product teams, and we'd typically have to do all this work to make it production-ready, even if the company wasn't ready to commit to investing long-term. So, the, like, the business equation was backwards. You'd, uh, incur this huge capital cost just to see if the product was worth investing in. And we still see that today with Fauna customers. Um, like, we have a lot of customers in the games industry, and they have the same problem. You don't know if the game is going to hit until you've deployed and marketed it. At that point, it's too late to do all this work to build scalable systems. So, you sort of do what you think is the minimum, hope for the best, um, and then try to clean up

But you forgo so much opportunity in that gap, and, like, there are cloud services that can help, because you can get hardware, but they still don't scale your actual application in the way that is technically possible. You have to make provisioning decisions. Usually the stuff that's more scalable supports primarily a key-value workload. Um, so you have to build a more sophisticated application. And just doing, having that experience over and over, and getting really good at it, but not being able to get ahead of the curve, is what led to Fauna. So, Fauna is designed to be that, that fabric, so to speak, that lets you model your data the way you want it to be modeled, and then run all these workloads safely together, small-scale, large-scale, cloud, on-premises, what have you. And it's totally doable. Like, information science doesn't prevent it

just takes a lot of work. Um, and most people are coming to, in particular, our market, from overly narrow viewpoints. Like, they'll focus on serving one query pattern really well, but nothing else. And that doesn't help you solve this problem, because you have to maintain, now, a half-dozen different systems, integrate them, and you lose all the benefits of having, um, monolithic database in the first place. So that's what Fauna is about. Um, there was a blog post just the other day from the Twitter team where they mentioned a bunch of the systems we built, uh, specifically that they're still using them, Flock and Haplo being the biggest ones. And that shows how, how hard these systems are to replace at scale, and how much path dependence there is in your infrastructure if you don't start from a flexible foundation. That's like eight years later

Eight years, yeah. Eight years of flow. Countless engineers who have come in and left and come in and left. And, and every six months I have at least one conversation about replacing Haplo, because that's technically a cache. And, uh, it always, uh, basically always ends up, you know, us punting the decision down the line. But we constantly have that examination. It's to replace. It's perfect

Brandon's not here. Brandon was the main Haplo engineer. Um, he's on our team as well at Fauna. Yeah, I think one of the funny things about being at Twitter is you have no, you have an endless supply of, you know, of like funny incidents and stories that you can tell, uh, which is actually quite helpful sometimes when you're pitching feces to talk about the crazy things that happened and how, you know, how they were solved by X, Y, and Z. Um, I think for us, you know, the, the stories that we draw inspiration from at Buoyant, uh, are really around kind of service to service communication and how that broke down in a variety of ways. There's so many ways that that can break and you don't really, I think, realize what's going to happen when you go into this world. Cause you have this model in your head of like, well, you know, how does my web browser talk to a web server? Okay. Like makes a request

And then, you know, if that request, uh, you know, it doesn't, if I don't get a response within 500 milliseconds or whatever, try it again. Like, try it three times and I just give up. Uh, and so you just apply that model to like, okay, this is how my services are going to talk to each other. And then what you find out is once you have like, you know, 10 services in a row that all have to coordinate a response, uh, that model breaks down in kind of quite horrible ways. And you have to start adding a little more intelligence to that. Um, so, you know, to give a couple examples, uh, we, we had multiple incidents at Twitter where a service would go down, uh, because someone made a mistake, right? You misconfigured something or you screwed up a deploy and the service went down. No problem. Let's try and bring it back up

And then you bring up one instance and all 3000 clients would immediately talk to it and it would fall over because there was too much load. You bring up another instance up, same thing would happen. So let's kill, you know, it's upstream dependencies, you know, so that we can bring this one up. Well, now we can't bring those up until we kind of follow this pattern all the way to the edge. So one isolated service, you know, even though you've gone through all this work of making this beautiful modular architecture and, you know, separate teams and whatever. Well, you still got all these operational dependencies where one service screwing up, uh, means that your whole site can fall down. Uh, and I think it's made worse by the fact that often, you know, load and, uh, failure are tightly related. Things fail under load

And so whenever you're trying to recover from that failure and you're adding more load, well, now you, you have, you end up with cascading failures or whatever they were, whatever they were called. It's actually, you know, it's a funny term, but it's actually kind of accurate. There are many ways for these failures to cascade. Um, another one that I really enjoy reminiscing on was, uh, we use Zookeeper as a source of service discovery. You know, this is the thing that tells you, you want to talk to service X, well, this tells you where all the instances of service X are. Uh, and someone went in and like, okay, Zookeeper goes down. Well, that's a problem. So let's make our, you know, let's make our, our systems resilient to that

So like, we'll, we'll cache some data. Um, and, but Zookeeper would go down and then sometimes it would come up and it would be empty. Or sometimes it would stay up and, you know, someone accidentally deleted all the Zookeeper data. So Zookeeper was up and responding. It was just saying, no, you have no services anywhere. And then of course, you know, all of Twitter would, would stop working. So being able to handle that kind of failure. So the, you know, the service discovery is up and it's responsive, but it's actually empty

Um, was something that also took time. You know, it's not the kind of thing that we really expected. Uh, it took time. And a lot of this logic eventually ended up being encoded in Finagle, which is Twitter's RPC library. Uh, in the buoyant world, uh, we have an open source project called Linkerd, which builds on top of Finagle. In some ways it's kind of like a thin wrapper on top of Finagle that, you know, lets you, uh, configure it with the YAML instead of by writing Scala code. Um, but we try and stay as, uh, as close to the Finagle code base as we can, because so many of those weird operational incidents have been encoded in, has been encoded in Finagle. So Finagle is very clever about how it talks to the Zookeeper or to any service discovery endpoint and how it caches that and, you know, how it, how it thinks about the health of an instance and what happens when, you know, Zookeeper or service discovery is telling the truth and what happens when it's lying

All that stuff gets encoded in Finagle and by proxy in Linkerd. So I think we're fortunate in that we get to kind of piggyback off of a lot of those kind of horrible, uh, uh, things that happened, encode them in Linkerd, and then, you know, you just use, you're using Linkerd, and it's like, uh, you know, I don't know, stuff is happening, but like the requests are still, are still flowing through. Um, so those are the, those incidents I think that come to mind. There's probably more for me. Uh, there were fun ones around like fires in the data center and rain in the data center. Like actual burning. Uh huh. Yep

Yep. Unfortunately not at the same time, because then the rain could at least put out the fire. Um, but, uh, those had, Finagle couldn't do a lot to help with those. Linkerd can't really protect you against fires in your data center, uh, except in the fire. Uh, except in certain situations. So you mentioned that, you know, like somebody would delete all the keepers. Is it the worst? What was the actual worst incident you can remember? Uh, worst, worst? Yeah. Gosh

The new ops person who rolled all the cache servers at once accidentally. Yeah. Oh, yeah, that was a good one. Was probably the best one. The black hole, the black hole replication table was pretty good too. Yeah, there was a newly hired ops person who was following a run book and ran the command specifically. And then the command specified that I think incorrectly stated roll all the cache servers instead of a single cache server. And that was really bad

Uh, wait, you were talking about, I think that particular instance was, I actually told him to roll all the caches with two minute gap in between. Remember, this is the time when every time you, you restart a cache, you throw about three, one to 5% whales. Like one to 5% of all the records would be whales, right? So we're like, so I, I've, I did that many times and I usually did it in the middle of the night because that's, that's where you get closer to the 1% and 5%. Um, so instead of, uh, doing one at a time, two minutes apart, I think that person just somehow confirmed, wrote a bash, you know, bash is hard, such that, uh, it basically did all of them at once with maybe a bit of weight at the end, which is a no-op. So, so, so, so, so, so, so, so, so, so, so, so, so, so, so, so, so, so, so, so, so, so, so, so, so, so, so, so, so, so, so, so, so, so, so, so, so, so, so, so, so, so, so, so, so, so, so, so, so, so, so, so, so, so, so, so, so, so, so, so, so, so, so, so, so, so, so, so, so, so, so, so, so, so, so, so, so, so, so, so, so, so, so, so, so, so, so, so, so, so, so, so, so, so, so, so, so, so, so, so, so, so, so, so, so, so, so, so, so, so, so That's why their systems suck. That's right. It is like a curious artifact of, like, Twitter in particular. They're, like, running out of data center capacity in 2010 forced all kinds of tradeoffs

It's like you just don't see from companies at that time. We were out the day I got there. And, like, it's a kids these days thing. But when I went to Nest and we were in Amazon, it's like I don't have to do any of the stuff I did at Twitter. Just buy more capacity. Like, efficiency, like whatever. You can just get more machines. Call up the Amazon people and we'll pay them 50 cents an hour or whatever

And it's far more efficient to just give Amazon money than it is to fix a system. Which is fine up until the point where at Nest we just couldn't get a bigger system. You're running on RDS and it's like this is the biggest one. Sorry. Maybe you should think about, like, Cassandra or something. Like, oh, but it's breaking now. Like, good luck to you. That was my first Christmas at Nest was that

Like, there are no more RDSs. And figure something out. Twitter had to do that early. But it's funny to think of what the infrastructure would have looked like if that NTTA thing hadn't been such a disaster early on. Like, if it had been possible to get more hardware. It was fine early on. But things had been. We were just behind the curve

Like, we had moved there from Joyant. Right. There were like a dozen Solaris boxes running over the end of my Australia. Like, NTTA was a big improvement. Another thing that's interesting. So, the way Twitter still operates cache now is that we don't have a single multi-tenant cluster. Like, a lot of databases are used. Which, for good reasons, right? Because cache is performance critical

You don't want someone's big-ass slow workload to slow your very agile and small queries. So, we want as much protection as possible. So, separating people into their own processes is a good idea. But as a result of that, if you, especially now, you know, Twitter is all, you know, microservices. So, every microservice will want to have their own cache. You end up having hundreds of them. So, it becomes a big burden for the team to support this many users. Not to mention the cost of creating and tracking who's using what

But every time anything goes wrong with anybody's service, they often go look at their chart and be like, oh, my tail latency for all my cache queries are pretty bad. Could be cache. Let me just ping the cache on call and get them to look at their service and make sure it's not. It's probably cache. Yeah. So, it's. They did that to Flock, too. Right

Flock was behind basically every query. So, if anything else was driving more load, the Flock graph would go up. And it's like, you know, look for your keys or the light's the best. So, this is true for cache. Actually, this is true for pretty much all the common database services in any, you know, large platform. But the problem with that is now every time there's an incident. There was a period of two years. Every time there was an incident, I'm in the room

Even though it had nothing to do with my service. It's just like a habit. Right. I tried very hard to sort of shake off that habit. And one of the best things you can do is to show people that your service is never slow. Well, never. Not technically. But you can get pretty close to that if you do all the things right

So, that's one of the things is I want my tail latencies to be really low even while I'm driving near red line kind of load. So, I have deniability. I can just tell people they fucked up themselves. That's actually the motivation for me to do a lot of things that I do is I do not want to be called up in the middle of the night to debug something that's not my problem. It's fun when you have one customer. You're just hanging out. It's not fun when you have 100 customers and everybody try to point fingers. Why don't you shift teams to one of the really, you know, services at the beginning part of the stack

The truth. Right. Exactly. Be up at the top and then all the problems come from, you know, deeper in the stack. That was my impression. But it's kind of an undersold, I think. It doesn't get talked about as much. But there was an arc in Twitter that followed a lot of the infrastructure changes around observability

And I think we went from Ganglia, which is terrible. What was after Ganglia? Oh, Ian's thing. Reunion. Yeah. Was there just for, like, collecting metrics and looking at stuff. And I think in Rails land at the time, it was pretty rough to get stuff out just as far as, like, latency and success rate and actual metrics. And I think as Twitter grew up, there was a huge investment in the observability stack to build out, like, Cuckoo and Viz on top of Cassandra, which ended up being a really, really nice system. And there was this kind of bending of the arc towards just, like, we will collect every metric possible and display it in Viz and Cuckoo, and that's going to be what we look at

And that was just too much cognitive load for people. Like, all the information's available, but you still can't tell if things are right. Like, I know that users are suffering for some reason, but I still can't pinpoint what's right. And then there was this kind of, like, settling back into there being pretty uniform dashboards across services. It's like, what is my latency? What's my success rate? If those two things are fine, then I probably don't need to look at anything else. If they're not fine, then maybe I should look at my providers and see if their latency and success rate is okay. it's probably me. If not, I can just sort of, like, pass along on call

And Streaming API got to do that sort of in a really crude way early on because we measured fan-out times. And so we'd get pinged because there were alerts on fan-out times. How long did it take a tweet to get to the Streaming API? But it was never really our fault. It was, like, all this stuff had to happen for a tweet to get found out, so we'd just be the canary in the coal mine. It's like, oh, not us. Cool, there's an incident. Thanks for waking me up. You should go page somebody else

And so getting that all tuned, especially as the org scaled and went from 80 engineers to 500 was really hard to get right. I think we spent a lot of time working on observability and, like, what to measure and what to alert on. And I think it really would have made engineering much, much more painful to work in had then not been improved. So I want to kind of, you know, ask you maybe now a question which I tried to solve several times. And I want to maybe through that question we'll kind of illuminate some pieces of kind of Twitter stack, right? So I did several meta-ops at Twitter. And my initial approach was, you know, you guys were in Twitter and you kind of discovered different pieces of Twitter, right? So you kind of form some kind of a stack, right? But it's a big organization, so you know, you know, maybe everyone kind of, you know, started early. And so you kind of have visibility. But, you know, later probably it grew, right? So every person essentially has a certain visibility in the organization

So I tried to kind of showcase this to SF Scala, right? Who majority of us did not work at Twitter, but we all use Twitter. So we all see tweets. And so my kind of Scala propaganda was very simple. You want to, you know, if you have doubts that Scala is fast, tweet about something. I'm going to like it and you will see the star appearing at your client very quickly, right? So the throughput of the whole thing, you know, is very fast, right? So I don't need to argue with people about Scala being slow anymore. I can just demonstrate that, you know, the whole cycle works and there is a lot of Scala inside. So I wanted to get several teams from Twitter to tell us how this life of a tweet actually works. And so my kind of naive approach was to get somebody from Finagle team, somebody, you know, customer facing some kind of API, right, the Finagle, then look at storage, right? It's stored somewhere

Maybe basically it will eventually go all the way to Oscar Boykin and data. And I wanted to see, you know, how money is made on this tweet, right? So somewhere during the life cycle, an ad is shown and Jax is, you know, big graph of money going up, right? So basically that was the ideal. So in between this plan and the actual meetup, all the talks became Finagle talks because Finagle team was in charge of this. And then I realized that a lot of people whom we got, you know, from Twitter are meeting each other or connecting at the meetup. So I wonder, right, is there anything about Twitter organizationally which makes, you know, basically what is the organizational features of Twitter which led to this stack as it is right now? Why was it so hard to put together people from Twitter to talk about the life of a tweet? And basically is it kind of technosocial feature of Twitter stack? Is it reflected in the way it's compartmentalized, right? And so kind of the follow-up question to this. So now three of you guys want to make businesses out of individual pieces. So you have probably an idea of your customers. And this idea is probably extrapolated from your internal customers or, you know, what people inside Twitter interacting with this piece are

So I just wonder what is your idea of your customers and how kind of, right, like I wonder, like there were some internal customers at Twitter of similar tools, right? And now you're targeting external customers in the world, right? I kind of want to, through that, I also want to understand maybe a little bit more how people kind of operate at Twitter, right, and how we can take this and bring it to the world. So I think the common theme at Twitter is these teams formed in a way that I think the promise was you can work on a thing and not have to worry about all kinds of other stuff. Like where I think in the early days with Rails, there was a huge amount of stuff you had to worry about if you were writing a front-end feature to get the feature shift. It was like MySQL stuff and Memcache stuff and Starling stuff and whatever to get a feature shift, demons. And as it evolved, it was just, no, I have a finagle client. I trust it does the right thing. I can call network stuff and it's going to work. And I'm going to serve it to some other queue and that's going to work

And as an engineer, I can just focus on doing my thing. And I think for various reasons, because Twitter was under stress from a number of angles, that promise never really materialized and you still had to worry about stuff. Like you still had to worry about the MySQL instances behind Flock or you still had to worry about the topology of cache servers. And you still had to worry about network RPCs sometimes. And so I think at least for me it was, but if you had a company that's just focused purely on solving this specific problem that Twitter tried to and got 75% of the way there, you could make things easier for a much broader range of people out there. And there were a bunch of those problems. We've chosen three. There's actually a couple of other Twitter-inspired infrastructure startups

Wayfront is basically taking the observability stack. Some of the other ones have already gotten acquired before they did anything. But I think especially Twitter had a unique combination of interesting problems to solve and these external pressures from data centers and other stuff that we were able to get 75% of the way there before more important things came along. And so there's always that. We could do it much better given more focus on the problem. One of the interesting things about Twitter from now that I've stepped out of it and kind of looking at it from the more corporate angle is how much of an investment in core infrastructure I'm making. You know, you think of it being this kind of toy-like, you know, UX where you're just like, ah, it's purposefully lightweight and easy to make a tweet. But, man, they poured, you know, developer year after developer year into systems like Mesos and Finagle and the database layer, whatever that stuff does down there

Like they really had a significant investment there. And I wonder if that at the time it felt right and it was good being there as like an infrastructure engineer, but I wonder if it was really warranted in the long run. Either way, I guess it was good for us. I think having later migrating to a microservice-oriented architecture means that the internal customers start to look a lot like external customers because they will look very different from each other. There is a very broad spectrum, at least from the services that I have contacted, in terms of their requirements and features and load. So basically, unlike a lot of companies that are even bigger than Twitter, who sometimes ended up building a very specialized type of infrastructure that is super cool but not necessarily applicable to anybody else. Twitter's angle is that it's just the high-level architecture decision may force this infrastructure teams to build solutions that are a little bit more generic. And that means the solutions are a little bit more usable for anybody, just random, right? So I definitely could see that in the system that I have to maintain

Do they call it microservices now? I don't think they had that word when I was there. I don't think they call it. They were still calling it SOA, which you can't say in the outside world because it's a dirty word. Yeah. And another thing is, so I think the value for someone coming from that kind of experience is a lot of systems are fine when you demo it. A lot of systems are fine when you run it for your beta, you know, even initial launch. The problem hits you hard when you go over that, you know, smooth initial ramp and then suddenly end up in the unknown territory. Being, using a solution by some people who's seen a little bit further along the curve means you can avoid hitting walls while being blind

I think that's a value. So if someone knows this thing runs on a thousand nodes and can't handle 10 million QPS, you are likely to be fine until you get there and that's a long way. So instead of making the same mistake and learning the lessons hard way, it's nice to sort of pass or spread that knowledge around and avoid. Basically it's time and resources wasted that can be used to produce something more interesting. I think some of the downsides there is I think as Twitter got driven more and more towards service-oriented architecture and microservices, service-oriented is like Tipco and stuff. Nobody wants that. You know, it got more limited, but in a way that was kind of liberating. It was like you didn't have a full computer at your disposal

You had Mesos. Like you weren't going to get storage. That was just going to go away. You didn't get to pick your ports. Like those got handed out to you. And in the beginning, I think as we were doing the initial migration over to Mesos, there was a lot of resistance. It's like, no, I'm a computer engineer. I deal with computers, not this container nonsense

But I think once that model got adopted and you were willing to forego some of the things like a durable file system, then your ability to work and not have to worry about a bunch of other stuff, like talking to provisioning people about getting physical machines, it was a reasonable trade off. There's a bunch of things you don't have to worry about. There's also a bunch of things you can't worry about because you just flat out can't do that. And I think that's kind of the difficulty in going out and selling Twitter stuff to other people, is you have to convince them that those things are worth giving up. And I think microservices and cloud-native stuff in general has that challenge. There's people who like live and die, bare metal, I have Ubuntu, and I have Chef, and that's how I roll. CoreOS seems crazy. Docker seems even crazier

Like, I'm not going to do this stuff. And you have to paint this picture of, but it got way better in Twitter when we were willing to make those trade offs and give up a bunch of stuff we could do with a promise that we didn't have to worry about a bunch of other stuff that we used to have to worry about. Yeah, I think that's true. It's certainly easier to pitch a post Twitter product to people who worked at Twitter. And it's not because we know them. It's because they're like, I can't believe I still have to deal with this nonsense. Yeah. I don't want to think about that anymore

I spent four years of my life worrying about the file system scheduler on my database. Like, why? You know, the cloud is here. Why is this still my problem? Yeah. People come out of Google and they're like, man, Google infrastructure is 20 years ahead of the rest of the world. I feel like coming out of Twitter, I was like, man, Twitter infrastructure was like 3.5 years ahead of the rest of the world. You know, and now I think especially with the explosion of Docker and Kubernetes and DCOS, like, we're really seeing a lot of the same and microservices, we're really seeing a lot of those same patterns play out, you know, in the rest of the world. It's still slow. It's still the details are different

You know, we didn't have Docker at Twitter, but we had C groups. It's, you know, kind of container-ish. We didn't have Kubernetes, but we had Mesos, you know, and we didn't have microservices. Well, we kind of did. We called it SOA. But, like, you see these patterns starting to emerge in the rest of the world. And so that's, I think that's gratifying to know, A, to know that, geez, we weren't crazy and we didn't go down this path that the rest of the world is never going to adopt. But also, you know, it's nice to have an idea in advance of the problems that people are going to run into

Right. What's going to happen when you run your database in this orchestrated containerized environment? What's going to happen when you have hundreds of services and, you know, you're relying on whatever crazy service discovery and you're just making regular old HTTP calls and praying for the best. Like, we have an idea of what's going to happen. And so I think the difference, like, yeah, I alluded to, is that your customers inside the company start to look more like real customers. And I think that the hallmarks there is that the difference, I did a meetup presentation the other day, that a lot of the premise was a lot of the advancements in tech, you know, whether it's Kubernetes or Docker or whatever, have really been smoothing out this relationship between ops and dev, which has been historically contentious. And you've created, I think, a really nice boundary for ops and dev to work together. It's like, look, dev, you don't have to worry about all this stuff. Like, you just don't get a computer

Sorry. You're going to get a stripped down abstraction of a computer. And that's going to be okay because we don't have to worry about, like, your OS and stuff. Twitter had a real fight deploying node to production, kind of sneaky-like for the build system. It's not going to run in production, but it runs on production machines. Well, yeah, but not, like, in production. And it was this huge hassle for ops to have to install node and try and operationalize it. You know, with Docker, if the interface is containers, they don't have to worry about that

But I think the flip side there is that there's not been that same sort of interface smoothing with customers and infrastructure. And I think as microservices get adopted more and more, you have internal people that are in your company who are going to behave more and more like customers. And I think the difference is you can't really negotiate with them at the same level you used to be able to with, like, dev and ops. Like, you went to the person who's had two desks down and said, hey, I need this thing on production machines. Can you do it? I think especially, I'm guessing, in Yao's case now at Twitter, like, there's a thousand engineers. You don't know the person who's asking for cache features. They don't want to have to know that they're actually asking for cache features. They just want this thing to work

I think customers have the same behavior. Like, there's not a negotiation. There's just an expectation you have to meet. And I think managing services in that way going forward, even internally, is going to be an interesting challenge, especially when there's hundreds of them within a company. So developer's job gets harder. I think manager's job gets harder. Where a manager used to manage, like, a monolithic piece of code, now they may manage 10 services deployed. And a director VP may be managing hundreds of services deployed

Like, just knowing what the hell is out there and who's using it and how they're doing becomes a huge challenge. Thank you. I'll ask one more question, and then you guys, please, ask yours. So this is of Scala, so I need to find a connection with Scala. And obviously, this is a joint with reactive microservices, which I think is the crux of the matter. But so I think this kind of ties into this customer question, right? So Twitter is the biggest call base in Scala, right? And the kind of biggest, largest amount of Scala developers. So I wonder how the fact that there is so many Scala developers you work with kind of affect your components. And when, you know, kind of other general, smarter than average customer now

Basically, you know, taking these pieces and put them in the, you know, wild, you know, outside world where, I wonder, like, do you see more Scala there? Do you have Scala customers? You know, basically, what kind of heritage, let's say, working with this big Scala code base left on your component? And kind of any interesting connection with Scala I'm curious about? So Fauna is implemented in Scala. The impact of Scala kind of ends at the Fauna protocol. It doesn't matter what language you use downstream in your application to consume the database. But we've had a good experience using Scala. When we set out to build this thing, we evaluated Go, Rust, Scala, C++, and Java. And we settled on Scala because we wanted the JVM for having a consistent deployment target, like a good operational story, especially in the enterprise and good performance. And we didn't believe that the Java type system would let us do what we wanted, which was to build an extremely sophisticated, essentially a monolithic process. Because Fauna is distributed, but there's just the one jar on each machine that does everything and manages all roles internally

And the internal interfaces are well encapsulated, but all in process. There's no service architecture there. And we wanted a type system that would let us get some safety out of something like that. We didn't believe Java or even C++ would let us push it that far. So we've had a good experience implementing the database in Scala. I think it's been very productive. And we have, it's pretty industry Scala. Like there's a monadic query evaluation system, but it's not like dependent types and Scala Z and that kind of thing going on

It's just like sophisticatedly typed. Is that a word? Sophisticatedly? Richly typed. Fancy. typed. Fancy, fancy. Richly typed Java, essentially. And then we get all the benefits of JVM and the familiarity operators and ourselves have with that and deploying across all kinds of crazy infrastructure. If the JVM is there, we're good

You can have whatever container thing going on. I don't know. You probably put it there. Yeah. Fatou. And it doesn't matter to us. And we can get the performance we need out of the system. In terms of customers, though, like our targets, like our adopters are working more in typical web and web services languages

So you get Python, Ruby, C Sharp, Java, and some Scala. But I think the sweet spot for Scala, at least the kind of Scala we build, is more systems level. And we're deliberately, like our product, the hook is that you don't have to do that stuff. So we put all the Scala behind the database interface and then all the monads and parallelization and magical performance happens there. And you can write a very complicated query, ship it as a single shot to the database and trust that it'll do the right thing. Yeah, for us, we were forced into it because we wanted to build on top of Finagle. And Finagle's in Scala. So we had no, you know, we had no, it was either rewrite everything, you know, in some other language or build on top of Finagle directly

And it was clear that was the thing that, you know, would save us a whole lot of time and energy. And then going forward, you know, Twitter is very nice to us. It production tests all of Finagle every day for us, you know, and make sure all those bugs get worked out. And once it's passed that process, well, we roll it up into Linkerd. And Linkerd, you know, we have a new Linkerd version. And look, everything's been super production tested at scale. It's interesting, I think, from the Linkerd perspective, because it's an open source project, but it's also an operational tool from the operational aspect. You don't have to know anything about Scala

In fact, we don't even use that word anywhere. We don't even use the JVM word. We're just like, hey, it's this process. You're probably running it inside a Docker container. It's like, you know, it does some stuff and, you know, you just do whatever you want to do. The flip side of that, being an open source project, is that we do accept contributions and we have people, you know, who want to do stuff. Or if you want particularly fancy behavior, well, you can't really express that in YAML. You have to write a plug-in

We've got a nice plug-in system by virtue of the JVM. But you're probably doing that in Scala. And so there we encounter, you know, a lot of people who don't know Scala who want to do something, and we try and guide them through that process. Honestly, it's a barrier. It's a, you know, it's a significantly rich language, and we have to spend a bunch of time helping people through that, which is not ideal, I think, at least by that metric, especially since the ecosystem that Linkerd operates in is very Go heavy, and Go for all of its many flaws is a very easy language to get started in. fact, I think it's almost the opposite of Scala, where Scala has a very expressive, you know, and beautiful language with kind of a horrible bill chain, and Go has a beautiful bill chain, but it's just like the worst, you know, it's, I won't say the worst, it's the least expressive, least beautiful language that one could possibly imagine, you know, without even the excuse of, you know, history to kind of explain away why it is the way it is. Anyways, so that part has not been ideal for us. On the other hand, our developers, you know, who by and large have had many years of professional Scala training at companies like Twitter, love working in Scala and allows them to do a lot of stuff in a way that they feel, you know, is safe and expressive and all that

So I'd say overall, it's a good relationship. I think the other maybe wrinkle for us is that since we're, since Linkerd is an operational tool, we care very much about its resource footprint. The JVM, you know, is not the place you would go if you really wanted to minimize resource footprint at the expense of everything else. So we do a lot of JVM tuning as well to minimize memory footprint, to reduce cores. And there, you know, we start getting into things like kind of the gory details of Netty and some of the lower level stuff where we have to avoid some of the Scala features in order to make it performant and in order to reduce the memory footprint. Some of that comes from finagle, honestly. And then finally, on, you know, point number seven on my list, really the code that we're writing is more finagle code than it is Scala code. Finagle has its own set of idioms and, you know, you write things in a very finagle-y way

So even if, you know, even if it were in a different language, I think there still would be that learning curve around finagle, honestly. But hopefully, for the most part, the Linkerd users are never really exposed to that. It's only if you really want to start contributing code or if you have a very sophisticated behavior that you need to encode as a plugin that you really get into these gory details. Cool. So I'll make two notes for our meetup. We need two courses, right? We need Scala for Linkerd developers, and we need Scala for finagle developers. Those two courses might be the same, almost the same. And this probably will have to involve 90 people, like, in performance, right? It will be JVM performance tuning, right? Yes

I mean, I agree with you. Similar to C++, Scala is a language toolkit more than a language. So you have to narrow the segment of it you want to use, build your own internal semantics, and DSLs, and then you build it to your domain, and then you can be incredibly productive. But there's a steep bar to getting to that point. And if you're like, I want to write a web app, let's use Scala. You're probably going to have a bad time compared to someone who can bash something out in Node or Ruby. And, you know, it's not correct. It's slow

It's not type safe. It probably doesn't scale. But they don't care because their web page is up. I think that's the number one lesson about Scala at Twitter. Like, I'm objectively observing because I actually don't write that much Scala. So I have no attachment, nothing. Me neither, honestly. But I tell people

One interesting other point is we actually have the same component written in Scala and Java and C++ for, don't ask me why, but that happened. And by far, writing a service in Scala on top of Finego basically was like 10x more productive than the others. I'm not even exaggerating. The C++ version took more than a year, and the Scala version was done in a month. I think this speaks a lot to the strength of Scala. But more importantly, and also the way that Twitter could build an infrastructure on top of Scala was it was very, very heavily regulated and templated. And I think that's really crucial, like you said, right? It's a toolkit. It's not – as a language, it's very flexible

So to basically use it in a more formal setting, you sort of have to have a dress code and spell out how you want it to look like. And if that's the case, then it's a very, very powerful language. Can you guys write a blog post how Scala performed all the others so we have like a nice use case? It's called a success story. That's a good anecdote, yeah. Well, the versions, they're not all open source, so it will be a little tricky. Also, the person who did the other versions are no longer working on the team. So I think at this point I would just be badmouthing them, right? So it's okay to badmouthing them in a very small setting. Let's not do that – take that to the written record

Who was it? I need a name. So we went kind of the opposite way of Boyan. Like I referenced my frustration with getting Scala Bills to work at Twitter as one of the big reasons I left Twitter. But for some reason went to Nest and did. You chose that beach to die on. I chose a very small beach. And all of a sudden it was, but look, there's so much more beach we can tackle at the same time. You should tackle the entire beach

I wanted to fix birdcage builds. It was fix Twitter builds. Like this is an impossible problem. It's like, yes, we know. Try. And then I left. So then I went to Nest, which is also like a heavy SBT user, but like the modern SBT. They'd made the mistake of contracting out a lot of their SBT build developments, and nobody internally understood what had happened

But there was a lot of magic in there, which really crippled us as we went forward. And so leaving Nest, I had, I like working in Scala, the language, and I just really didn't like working with Scala, the build chain, anymore. So we evaluated almost exactly the same things. Like Rust was too early. Java 8 was coming out, but it was also like JVM-centric. And for our target, we were going to build on top of Nginx with the module written in C. That was kind of a choice we'd made. The rest of the stuff we had more options around were really small CLIs, which isn't really the JVM sweet spot and definitely not Scala sweet spot

We planned to open source those. And so that was another thing. Like we're going to open source small CLIs to ops people. The JVM seems like a tough sell there. And so we went with Go, which seemed like a reasonable sell. As long as we're writing the CLIs in Go, we have the small web service to write. Instead of introducing a new language, maybe we should do that in Go. And that's been, I think, the least sort of positive choice we made, the least good choice

The CLIs in Go, it's great. Like you get a good CLI wrapper. You get all kinds of good systems integration. Those things are really small. It's great. Doing a web app where you want things like patterns for dealing with JSON and getting that talking to a database, it kind of naturally leads you to, it would be cool if we had generics, but we don't. You have four loops. You have two kinds of equals

You have pointers if you think about it, but don't. And so we end up doing cogeneration instead, which is its own, like you have, you can have all the generics you want as long as you script them up yourself. So that's cool. And so I think if we had that to do over again, that's maybe the one choice that we'd undo and think about doing something based on the JVM instead, maybe. That said, the tool chain is fantastic. And so like my one sort of, you know, disappointment looking back at the Scala ecosystem is that there wasn't a bigger investment in making things like SBT and Zinc more available and more performant to get build times down. Because it's just sitting in a Go compile loop where it's just seconds to get my thing built and the compiler tells me if it's wrong and I can go about my day is super cool. Of course, starting with SBT and waiting five seconds for it to compile the build file and get to a REPL where I can then start my build and then wait longer for it to do stuff is just now I'm reading Twitter and rage tweeting

And after I get done with that, then I'll come back and check my build. Usually that's a half hour later. Because the tweeting, not the build. Yeah, it's sort of like if your program's going to, take over the whole machine, use Scala. But you've got to realize that that's what Scala is. So like SBT will also take over your whole machine. You didn't really want that, but that's where you're at. Yes

And we use SBT. It's like kind of okay. We're just building the one jar ultimately. So we don't have to get too crazy with plugins and dependency trees and stuff. But you just try not to get too weird. The one thing that I think we've been much more pleasantly surprised with is Go's vendoring systems just seems to be much easier for us to deal with. We didn't cover it here, but Twitter also went through a ton of like code repository gymnastics from this sort of like anarcho-syndicalist, like my syndicate has their own. I'm a big fan

From a code organization standpoint, it's like our syndicate has their own repo and you can't get into it. It's like we'll just give you artifacts out of it. And that optimized for some things at the detriment of others. And Twitter bent back towards this monorepo, which I vigorously fought. And now I'm a big fan of after going to other places. It just optimizes for different things. We're in a monorepo just because we're small. We have a bunch of things that are logically bound together and it's nice to be able to do a single build

And when we open source stuff, we don't want to open source a monorepo because that's just mean. Google, it's interesting to look at Google. And I think one of the reasons Google does very little open source is because they were monorepo from the start. And so things like stubby and chubby have dependencies all the way down to proprietary stuff that are impossible to extract in open source. Twitter made different decisions for better or worse in the early days. But I think that open source predilection carried through goes vendoring system as opposed to the horror show that is Ivy. Makes that kind of moving things around much, much easier. I'm really surprised the Java ecosystem hasn't come up with something better than Maven and Ivy at this point

But so it goes. Maybe we'll open the floor to the audience. Questions? So while you're guys thinking about questions, I just now realized, right, because of SBT, we have so many good tweets from inside Twitter. I just realized, right, this feels like that. That's why there is so... Right, that was the secret to Twitter's success. Yes. Those developer SBT-driven tweets

And a lot of communication. No wonder my tweets are terrible. I'm, like, GCC is just too fast. You need to have a sleep in there. Sleep 30. Make is nice. Yes. Like, I don't..

That's arguable. Compared to everything else out there. Make is okay. Yeah, we have recursive make files, and it works just fine. The really interesting thing for us doing, like, a full stack deal is it is amazing how much better JavaScript development has gotten. Our JavaScript code, our front-end code is much, much more reactive and functional than our back-end code, which is, I think, really disappoints our server engineers. But if you've done things like React and Redux, it's a lot of the same functional programming concepts. I think even though we do know Scala today, all of our engineers, all of our systems engineers came from Twitter and are heavily experienced with Scala

And I think there's no question that's made them better engineers just by being exposed to different approaches. And even being able to look at Go, even though it doesn't have generics, it still has functions as primitives, which allows you to compose things in ways that are, I think, much more flexible than what you'd have if you just started out with, you know, Java 6 that didn't have that. Yeah, Fauna exposes a very functional, mostly immutable query language. And it's still a version of the relational algebra fundamentally, but pushing those kind of Scala concepts, like, that has affected our customers. And they like working with a much cleaner, less procedural interface to the database. And that gives us the ability to parallelize everything on our side because there's not implicit dependencies we have to guard against and that kind of thing. So you map and fold and for each and join and do all your usual stuff, but you're doing it against a distributed pool of data, not just your local heap. And you can do that in any language

Some are a little more direct to implement the DSL than others, but it always works. So they do get, our customers do get to experience, like, I mean, arguably the best part of Scala as exposed through the database. I think it was something Twitter invested in, not intentionally, but in the early days, like Steve and Nick and Marius and trying to remember other people. We're pretty good about trying to evangelize functional approaches to things without being Scala specific because the whole company didn't do Scala. But, like, people were super patient and nice and not assholes. And, like, look, you know, fold left is a pretty cool concept in general that you can apply to Ruby because it also has functions. And you can apply to JavaScript because it also has functions. And these are ways that, like, help you think about systems and design better things as reference to these other papers from the 60s or whatever

It was, like, pre-papers we love, but I think a pretty, like, one of the best things about Twitter 2009, 2010 was people just wanted to get better at what they did. And this has been, like, in contrast to other companies I've worked at where people just wanted to come get a job done and go home, which is totally fine. But being part of that environment where people just legitimately wanted to get better at what they were doing was really, really great to work at. Yeah, I think Twitter was lucky in that it had a couple senior engineers who were very influential who also managed to walk the line between wanting performance and kind of realistic behaviors as well as elegance and, you know, expressiveness and didn't go too far down in either direction to the detriment of the other. I think we got really lucky in that respect. I think we have an audience question. Yep. Please introduce yourself and ask a question

My name is Gabriel. I'm at a company called Blender Workshop. And we do robots, but we have a small cloud backend that's playing on a scale of microservices. I inherited some of that, and I'm curious, because you guys have done this SOA on Twitter and watched that grow, and now kind of the whole microservices thing is very popular outside, you've expressed how the SOA makes sense for the larger organization like Twitter. Is there kind of like an organization size where you think it doesn't make sense to use kind of like a small microservice architecture, or do you think it works all the way down? So, uh, the machine doesn't care where your service boundaries are. Like, in the end, you're writing one end-to-end program. So, in my mind, and, like, what I did at Twitter, and it doesn't really affect us here, because the monolith ships sail up long ago, was, like, you draw your service boundaries where the human communication boundaries should be. It really has nothing to do with the architecture

And ultimately, a monolithic process is usually more optimal, because you can share more information, which is required for, um, meeting, or improving performance. So I think, like, you know, how big is your team? If you have two teams working independently, and they're stepping on each other's toes, then that's probably a good place to think about drawing a service boundary, if you can find a boundary that's not going to shift very much in the foreseeable future. And obviously, like, you know, front-end versus back-end is historically, and the database behind that, those historically have been the boundaries. And then, as, like, all these web-scale companies started wildly exceeding what were, at the time, typically just, like, 20 or 30 people working on these end-to-end web systems, started to run into these same problems. And then, I think, like, Twitter, in particular, started as a monolith, broke into core services, then went overboard into microservices, and now has, like, come back trying to find where the right balance is. But it's a human negotiation. There's not a technical answer. I think one of the dangers of, like, I think the stat from Uber was they have several thousand services in production

It's, like, four per engineer. Yeah. And I think the danger there is, like, who owns those services. And, like, number one, it shouldn't be a single person. That's bad. It should be a team. But if a team is two, let's be super generous, then every team owns eight services. That seems, unless you have fantastic tooling for that, which is a decision you can make, that seems like overkill

And so I think the balance really is, like Evan said, around those human interactions, it's great. Twitter had a couple of cases where the human interactions weren't just who's working on the thing, but when you have a deploy cycle for the thing. Like, I think auth got split out when we were doing the upgrade from OAuth 1 to 2 or some other, like, crazy requirement we had in auth, that it was impossible to get new auth stuff deployed through the monorail because deploys were so backed up. I think there were also some efficiency concerns. Like, Password was a weird early one. It was more, I think, operational with the number of people working on it because it was only one. I mean, deployment was a communication bound. Yeah, yeah, it's true

Fundamentally. There is another side that's not mentioned too much when you split things into different services. Like, that's the hard truth of distributed systems. Once you move out of a single address space, a lot of things become untraceable by nature, right, because not everything is asynchronous. You have to do a lot of investment, which even Twitter cannot keep up with, to keep that sort of debuggability and observability in the platform. And actually, that's the one thing that people, why do people come to infrastructure services? Because they don't know the causal relationship, right? And that's fundamentally hard in distributed systems. So I think, personally, I think the natural place to start with, because you inherit something, that's a different story. But the natural place to start is start with a monolith because everything is easy to reason

And if you are too big or if different parts just have way different performance characteristics and need to scale independently for the best efficiency, then you start strategically break them out. It's basically not for the sake of having multiple services. And I think that's, like, one of our core premises as a company is that's never going to be exactly right for you. Like, you're always going to want to tweak something, and it's scary. And having tools that help you, you know, sort of modulate your architecture in a way that's not super risky lets you constantly keep that closer to optimal. That's a lot. One question there. So Twitter has built a lot of things from scratch, right? Kind of, kind of, or or Manhattan, or Kestrel, or Pence, a lot of things

Maybe there is less of that now, these days. But the problem is that in 2017, we're still stuck with this technical chaser that you guys made, in fact, seven years ago. The question is, do you think Living Twitter has overinvested at some point in things building it from scratch versus taking something off the shelf? Was there a chance to take something off the shelf? There wasn't much off the shelf when a lot of those things came into existence. Yeah, nothing worked. at all. Like, as an example, like, the proxy Twitter built from scratch was built from scratch because at the time, I think Nginx's documentation had not been translated to English. Yeah, that sounds right. So we had very few C programmers

We had zero Russian-speaking C programmers. So, and I think, like, the experiments there were that, like, Nginx worked, and as we're using it, it was like, this seems to work pretty well. But the reasons for doing something from scratch at that point were we wanted to use a common networking stack across a bunch of things, and Scala was a better choice for that. And also just, I think Twitter's story was there were things that weren't quite there. And so MySQL we used heavily. Like, we didn't write our own database from the beginning. Manhattan was a response to a very specific storage problem that, I think, it was, like, the third response to a very specific story. That was actually my time, yeah

It fixed, I think, a storage problem that was there that there were two false starts on. Kestrel, like, you were there for the original Kestrel, like, studies, but I think they studied a lot of, like, RabbitMQ and NQP and, like, other things that were out at the time. And Twitter's workloads were just weird. Like, most of those things were written for investment banks. It was high-frequency trading stuff. Like, Tibco was there. Sonic was there. Rabbit was fairly new

Yeah, it's before Kestrel. There is a... Starlight. But there's a Ruby version. But there was also OpenFire and Rabbit. Neither of them were good. Like, at the time, like, Twitter's problem was pure scale. Like, now you can get off-the-shelf stuff that will scale, and all your problems are adjacent to scale, like isolation and high availability and Mali Data Center and security

But, like, you can get some shit from the cloud that, like, has some scale. You couldn't do that at the time. And stuff that was scalable would be like, oh, we did, like, 1,000 RPS on a single node. You want two nodes? That's crazy talk. But this one dev, they have an experimental branch. Use that. Of course it doesn't work. Nothing worked

Like, literally nothing worked. And we tried all this stuff. Like, we looked at the clustering systems for relational databases. They were all horribly and fundamentally broken. We looked at all these queuing systems. They were designed for very low message velocities. Their durability guarantees were false. They just didn't do what they said on the tin

And we tried Cassandra at the time and encouraged other people in the startup community to invest in it, too. But we couldn't get it over the finish line for operational data only for time series. And that was, like, that was, like, that was the thing that was, like, you know, talked up on the conference circuits that was supposed to be the best system. Like, when Facebook open sourced it, it literally didn't build. So that was the quality of, like, in particular open source infrastructure at the time. You had, like, the Perl stack, the Rails stack, and then, like, random academic quality Java tools. It was before Hadoop. Like, there was nothing

It was thrift. We used thrift. We didn't invent our own, you know. That didn't work either. That didn't do. That clearly didn't be a little bit. That didn't work, but we used it. So I think it's, you don't see it being talked about in blog posts and stuff, but for every project that is invented, right, there was, at least most of the time, there was investigation into what's off the shelf, and let's throw it into our production or staging and see how well it does

Oh, actually, it's terribly wrong or broken, so maybe we should do our own. I think that conversation pretty much always happened. It would be hugely irresponsible if it didn't, right? Another thing is it's easy to have a lot of marketing fluff if you don't have to look at the details and if you don't have to test it very hard. And most of the benchmarks I've seen are just all poorly designed and sometimes intentionally misleading. Those things you don't discover until you have to actually handle the load. I'm not saying Twitter has not overinvested in their own solutions. But generally, there is a good reason for that. And going forward, if the quality of open source solutions improve, or if people like these guys put out better solutions, just generic solutions on the market, then it should, as the quality of available tools improve, you should see fewer projects starting from within

Kids these days have it easy. They've got Pelican. They've got Fauna. They do Fauna. I think even with the developer tools, like you mentioned Pants, but as somebody experienced in evaluating a ton of build systems at the time, Twitter went from Ant. Pants is, I think, little known. The original term was Python Ants. So Python front-end for Ant build files was the original thing

It's like CMake. We looked at Make. I had a thing that translated SBT build files to Make files, and it was awesome. Yeah, it was good. It was awesome. I liked it. The downside is there's no command line test runner for JUnit, and I didn't want to build one, so I gave up. But no op builds were like 10 milliseconds

It was amazing. We looked at SBT 07. We looked at SBT 10, which was basically not SBT at all. And I think for the kind of development Twitter wanted to do, which was a monorepo with fast builds for teams that needed to build a sparse tree, like SBT just wasn't designed for that. And nothing on the market was. Yeah, Scala itself barely worked at the time. It was super buggy. And it caused a lot of internal dissension and infighting

Like, is it production? Like, is it really worth investing in or not? Some people wanted to use Java, but, you know, Java doesn't. Like, Java kind of sucks as a language. People coming from Ruby in particular were not interested in going back to C land and only having procedural code and that kind of thing. I mean, I think, like, the components we used that stuck were NLDB, Memcached, and Redis from the days when Redis was a single C file. Like, everything else got, like, completely rewritten. Even, like, we rewrote the Ruby garbage collector three times. Yeah, I think even with the build system, the reference was always Google. And I think at the time, Google had a team of, like, 40 people working on Blaze

Like, that was Google's developer infrastructure team was, like, 40 people. And Twitter was having, like, a big gut-wrenching decision about whether to commit two. There were zero full-time people working on builds. SBT was... 40% of a person. Yeah. SBT was one person spare time. And that was the state of the art of Scala Build Tools

Like, Ant, you know, all these things. Gradle. These are all, like, you know, passion projects for people. But if you want to build a build system that works at scale, you'd have the ex-Googlers that's, like, just give me Blaze. It's like, all right, give us 40 people for several years to go build Blaze, which isn't just a build system. It's a distributed system that happens to have, you know, object files as output. But there's, like, a distributed cache. There's all kinds of other crazy stuff that goes into that

And, you know, there's just no systems like that at the time. I think even now there's no build systems that really you can drop into a company with 1,000 engineers and expect to be awesome. There's other systems, but build systems are still, like, universally painful at that scale. All right, audience, what else you got? All the questions? And all of them? So some of you run the infrastructure startups, and we know some examples of successful infrastructure startups, like maybe GitHub or MongoDB, and some of them failed, like RethinkDB. What do you think is the secret to maintain the successful infrastructure startup or open source startup that's part of it? Okay, read the RethinkDB postmortem out. Yeah, definitely don't be a database startup, whatever you do. That's just, yeah. You're not wrong

I mean, you should ask us that in a couple years. But I think just you've got to build something you can sell. And you've got to have that in mind in the early days. It's like, there are companies that are, you know, you can get adoption through open source, especially if the thing's easy to drop in, but that also makes it easy to rip out. And then, like, for Rethink, like, you know, they didn't bring any products you could pay for to market until, like, literally months before they ran out of money. And things have changed, too. Like, people don't want to operate open source in the same way they did. They want to use cloud services and serverless stuff

And, like, if you can push the ease of use and ease of adoption barrier farther than anyone else, then you'll get users. It just takes time. But when you don't, and then you try to, like, get adoption from one community, then convert it into sales to another community, you can't bridge that gap. So you have to consider sort of the entire ecosystem around your product from the beginning. And, I mean, even, like, with you getting adoption for Pelican Cash, like, and even internally, you have to do all these business stuff, like, market it properly, understand the initial use cases that get people hooked, and then show them the garden path to the full power of the system and all that kind of thing. Yeah, I think people's appetite for being systems engineers is just diminishing. Like, I was at EMC before Twitter, and the big reason I left was just working on data center systems management stuff. And I was like, hmm, I wonder how many data centers there's going to be in a decade

Like, how many companies are going to run their own DC? And the answer is, like, not as many. So glad I got out of there. But then I went to be a systems engineer at Twitter, and it didn't make the conscious leap at the time. But it's like, hmm, how many people are going to actually write their own systems infrastructure in a decade? And probably not as many. And so I think, you know, for us, if you can give people either a new capability that they didn't have without your stuff, or allow them to not worry about something they have to worry about today, then fine, you can charge them money for that. The nice thing is people are used to paying. Like, Twitter wouldn't pay for shit. Like, if you wanted Twitter to buy a piece of software in 2009, it was super tough

Well, because it didn't work. Well. I think it was more than that, though. I think it was more than that. Yeah, but people just weren't coached to, like, especially in, like, San Francisco web startup land, people weren't coached to buy stuff. And people in EMC land, it was like steak dinners and handshakes between two gentlemen who owned boats to foist some piece of software on you that you were going to pay, you know, seven figures for in perpetuity for support. I think Amazon's really changed the game. People are used to the fact that they write a check at the end of every month for services rendered

That's cool. We can be some of those services rendered. It doesn't all come from AWS. And so I think that's the nice part. If you provide value, people are willing to pay for it. It's not a constant, like, but there's an open source thing I can do for free because I'm willing to worry about that stuff because I don't want to pay. And I think that calculus has changed quite a bit. The flip side is you better provide something that's one click and everything's done

Yeah. When they do have to worry about it, like, the social contract changes. Not social contract, I guess. Financial contract. The actual contract. The promise is, like, you told me I wouldn't have to worry about this, and it's broke. And now I've given you money, and now I'm super upset. Whereas open source was, like, a sternly worded message on the mailing list was, like, it gets different

Yeah. I think one of the things that took me a little bit of time to get my head around moving from being an engineer and an open source user to being now a purveyor. You know, I'm now the guy in the alley with, like, the trench coat full of open source stuff. And I'm like, come on, first time's free. You were always that guy. Okay. All right. You got me

But companies are always paying, right? They're always paying for stuff. They're either paying, you know, their engineers to build something, or they're going to pay for, you know, they're going to pay another company to provide it for them. And so they're used to paying for money. They're going to be paying no matter what. So once you enter that equation, once you understand that, you know, they're going to be paying one way or the other, the nice thing about being in the infrastructure world, you know, which is selling stuff to companies by its nature. You're not selling infrastructure to a consumer. You're selling it to a company. Is you just have to make the equation

You just have to make the value proposition. Like, look, you're going to be paying way more if you don't buy this than if you do. All right. So it's a very different world from, I think, having to convince someone to spend 99 cents for this iTunes app. You know, you've got someone, you've got a company that's paying money no matter what. You can convince them to shift where that stream of money goes and then to stand underneath it. Yeah. It's not to demoralize, I'm assuming most people in the audience are scholar engineers, which lends itself not so much to front-end development

Any front-end developers? Thanks. But, I mean, for most people working as systems engineers here, like, your company's core business is not building APIs. Like, your company's core business is delivering some other experience to customers. Yeah, I mean, I just wanted to keep the website up. Right. Like, there was no more, like, the entire time, it's just, like, make throughput better. Right. That's the actual goal

And so I think, like, going to... It doesn't matter what architecture is going on. Yeah, going to a place like Nest, it was made very, very clear that what I was doing was a cost. It was not the core purpose of that business. I was a critical component. But if they could just, like, spend 50 bucks a month and replace me, like, that was the brutal calculus. And so I think that that's going to be, like, the trajectory of a lot of web companies is, like, infrastructure is a tax. And maybe I'm willing to pay to deliver some, like, specific product feature

But the more of that, I can just outsource. Rule number one of business, if you have a problem, make it somebody else's problem. I think that's going to be the path going forward. So if we can be the other person whose problem it becomes in exchange for currency, then that's good. So your products are putting hardworking Scala engineers out of jobs today. My product is allowing hardworking Scala engineers to do more valuable stuff. Maybe that's writing JavaScript. I don't know

Other questions? Well, then, let's thank the panelists. Thank you very much, guys. Thank you.