Scale By The Bay 2019: Kaoru Kohashigawa & Tabitha Blagdon, Scaling Financial Automation on TypeBus
[Music] right good afternoon everybody you guys enjoying the conference yeah awesome so my name is Tabatha I call a car over with tally technologies and we are a consumer fin tech company based here in San Francisco and we build consumer apps to help people with financial automation so I'm really excited to be here today and share with you a little bit about our story and our journey and how our systems have evolved over time and how we've really been able to add some innovation and fund into building financial assistance in general how many of you have worked in financial services or a fin tech company before cool awesome anyone has anyone here worked at a earlier stage startup maybe serious here below etc awesome great so it's it's gonna be great it's exciting as all of you know in that space and and so I am very lucky to have been at tali for about two years and in that time I've seen our systems and our technologies and how we build things evolve and so what this presentation will go over is to share me a little bit about our journey and how we've transitioned from where we used to be to a new framework called typists which is how we build reactive micro services I'm over Kafka using aqua streams an aqua cluster so gonna be great so today we're gonna talk about four things first I just want to tell you a little bit more about tally and our story the problems that we're trying to solve as a company so you have a good understanding of what we're trying to do and then why we're building the products that we're building the second thing I want to go into is that journey to type us like what were those triggers what were those things that led us to evolve our technologies that led us to want to all the changing challenges he went we went through to be able to want us to involve to a different system if you will kuru over here is gonna talk about type us in the implementation in detail so he works on the team that maintains the framework and production eise's that framework so he'd be able to go deep into detail there and then we'll close off and let you know about some of the things were working on going forward good yes okay all right so I so when I was doing my search for what I wanted to go next I knew I wanted to go in FinTech I loved finance and the blend of finance and technology and helping people with their personal finance and what stuck out to me about tally and why I came over here was our mission our mission to make people feel less stressed and better off financially there's not a lot of companies that I interviewed at and looked at where that less stress piece that emotional piece of helping people work their finances it's really a big part of what they do and it's something that we believe in and so Itali we truly believe in complete financial automation for the average American we want to build a system an intelligence service in the background and ambient service that knows who you are what you do what your financial goals are and it does a lot of the thinking for you when it comes to your finances and it also doesn't work for you automatically if here's what we found is a stress of financial management especially paying off debt is crippling for a lot of Americans and that's something that we've been wanting to solve and so people can spend their time focusing on spending time with their families anytime with their kids or just their careers so when you look at just where we want to prioritize helping people there's one trillion dollars worth of credit card debt in the United States it's actually been growing quite a bit and 44 percent of households have credit card debts that are $15,000 or more and as you know credit card interest rates are very expensive right so it's a crippling level of stress and anxiety that we found in fact if you look at people with good credit scores they don't even have bad credit score just good credit stores alone 2/3 of them actually had medical levels of anxiety when it comes to managing their debt and trying to overcome it and when I say medical it means they're seeking professional help and so this is this speaks to us like we we knew we wanted to at least start with this financial job and help people with that and so the first product that we launched is what we call tally cards you want just about two and a half years ago it took a while to build but we finally launched a two and a half years ago and what it does is it helps people manage a credit card debt so our users sign up they upload their multiple credit cards to our systems we check their credit score and we offer them an automated a line of credit and then we use that and optimally pay off their credit cards for them and they only have to pay us one time a line of credit that we offer them is lower interest rate than what their credit card paying and so we've been able to help a lot of our users pay off their credit card debt faster sometimes 10 to 12 years faster and save thousand dollars an interest just by doing this simple job for them and so the job I mean the app is doing really well we've helped a lot of people but if you remember our mission which is to help people become less stressed and better off financially this is only one subset of people in the u.s. I need some help right we underwrite our own loans so we can only because of the amount of risk we can take right now we cannot offer loans to people with a FICO score of lower than 660 so of course that funnel of people we can help is a lot smaller and so we knew you know we need to take financial and automation and scale it to the millions right and so where we want to help next and what the product we're building next and which is the impetus of why we're moving to this different framework by the way is we want to help people with savings it is a top of the funnel if you can't save you can't set aside money you cannot pay off your debt so over the last year my team and I have been working on a called tally save or in private beta it hasn't fully been launched it'll be launched next year but it's a first app that rewards you for saving money and it's all for free we found people want to save they just don't do it for a lot of reasons and so we want to make saving easy we want to make saving delightful rewarding and we want people to see their progress and so MVP version of the app is it's pretty straightforward you can save as little as automatic as five dollars a week we reward you and we give you points for each time you do something good for your finances you can take those points and you can redeem credit cards I'm not credit cards but gift cards you can redeem you can pay off you can contribute to a charity and you can help save for different goals pay off your credit card address so that is that is the background on what is driving us and moving us to a new technological framework so unregenerate here at tallies aware scholar shop and we've we're excited like we're getting this new app but we need a build over the last year and the requirements to build this app was as follows first of all product told us we need to build this app and have it have the potential to scale to the millions which is it has to have that room the second thing we need to build is that reward system right we offer real time points we award points and that has to be really responsive our systems have to respond and react and process events in real time I'm also the money transfer process it has to be reliable it has to be something that is consistent reliable and it because it's a key user function that we have and so we knew that those are things that we need to build for our constraints is we want to stay within the Scylla ecosystem we're passionate about Scala and functional program in general and we also wanted to avoid some of the engineering challenges that we went through in the early stages of our company as we're trying to build products really fast so I want to start off and go over what some of those challenges were so you can understand why we move to type s so in the past and our tally cars product actually still runs on this system we build services pretty traditionally you know five six years ago this is how most services were built right so client talks to the elbe we go to goes into the API layer we use akka HTTP and then down to the server's service layer so IDL heal thrift finagle etc all the way down to repository to the database and you can see there was schema first there many communication layers that communications would have to go through a top of that multiple layers of serialization and deserialization a lot of boilerplate code gen etc services talk to each other through our PC and so we had to pull in dependencies from the clients of each service etc so it it wasn't easy to see the bus service I mean it worked fine for us but it was very error-prone there's a lot of areas that people can make mistakes and I know some of you might have worked in this area before so it works really well for us it's just we knew we needed something different for Cali save another challenge that we had was that we started building what was called a distributed monolith so even though they were microservices they were so intertwined that they really weren't right so tight coupling complex web of dependencies we had shared data stores our code was shared or data models were shared and pulled in the responsiveness resiliency of those systems was not that great at top of that our data pipeline was really brittle they were consuming they were pretty much each yelling data from our databases nothing was a real-time eventing wasn't really a first-class citizen we didn't really use a gun t and so that was exactly where you wanted to move to so when you add all of these problems together what you have is just really slow error-prone development cycles that was frustrating for engineers I mean onboarding difficult because the systems are complex and hard to reason about and we knew if we wanted to scale not just for a product of our organization we had to figure out a different way to build things so tally save when we start building a new system we wanted to take a risk and try some new we want to stay within Scala but try to do things differently and go reactive and make and build an architecture event based so what we did was we made a wish list of things that we needed for tally safe right we knew that we wanted the Avenger of an architecture to be the center of what we build we wanted to make sure that if we can build a framer that can help us build these reactant maker services I wanted we wanted to have an API that was easy for our developers to be able to hook onto Kafka hook on to whatever transport layering we're gonna use and people to speed up services easily we wanted the ability for code code first data modeling to prevent some of those serialization issues that we saw before wanted to avoid traditional service discovery and also we wanted to make a synchronous jobs and managing throughput for those asynchronous job losses the money transfer jobs pretty easy so it's a hefty wish list but we wanted to build a framer that can help us with this when you were looking around and this was you know a year and a half ago or so it was hard for us to find something that truly fit all these needs we had an engineer a really smart guy on in Vancouver and he was working on a side project that he called typos it was a very light version of this it didn't cover all the features but it was something he brought to us and he suggested we try looking at this as a framework and tool kit to build micro-services and we're like you know why not why not try it so we took a look at it and we really liked it we thought there's a lot of potential and so we worked together and production eyes I maintain that codebase and we now use type us to build micro services for tally safe and so what type us is maybe after tally touches that you can call it tally bus but it is a mildly opinionated framework for building reactive micro services in it supports and abstracts away a request and response pattern over a transport protocol so we used Kafka right now but it could really lay over anything it has a easy-to-use API so it's easy for developers to interact and just either consume data of Kafka etc and also spin up services and service contracts and so it's solved a lot of problems for us we've worked a lot and just to give you a sense of at a high level how our services work on this platform Peter birch services is built in this type of framework and they all communicate and interact with each other through Kafka so the eventbus is really the source of truth so that is at a high level is how our evolution of technologies has come about and now I'll pass it on to Carew and he will go over in more detail about how type s works under the hood great Thank You tabatha all right so let's talk about type us implementation so just the disclaimer here I'm gonna go over some concepts from Kafka hakka streams and aqua clusters I'm not gonna get a chance to really dive deep into the details of these but I am gonna be able to show you how we're combining these technologies to hit our design goals for typos so let's bring back those these angles the first thing we want to do is have convenient service updates so what does this mean as a service I want to be able to publish my state to the world and have anybody listen to that update by the very virtue of using typos we can accomplish this or sorry Kafka so I can publish my state to Kafka and other services can know that I mean it's a change cool we're done right let's just go home no I'm kidding all right so let's bring back those design goals we met the first one the second one we want to do is have an easy and convenient API so what does this really mean is that our engineers should be able to publish consume and also do requests in response patterns with a single API all right so what does that look like let's bring up some code here we have a cat client in this cat client allows you to get a cat given a name so we have a method definition here on line two that's give me a cab by the name it's going to return a future once that future resolves you'll get back a cat and then you go on your merry way on line six what we're doing is we're telling the Kafka consumer to look out for a cat message I'll get into a little bit more detail on how we're accomplishing that on the server side code we have a cat ranch and so this cat ranch is looking out for fetch cat messages and the fetch cat once it has a fetch cat message it'll pass into an eternal actor and then send back a cat to type bus and the reason why we're returning a future is so that type bus knows that the message handler has successfully handled the message one side knows that then it can confirm it back with Kafka on line eight we have a listen to adopt cat message and this is basically the SUBSCRIBE pattern so if I want to know that you adopted a cat that you adopted a cat or that you adopted a cat I can listen to those events do something with that and then confirm it back to Kafka to make sure that I have successfully processed that message line 12 and 13 is what we're doing is we're registering these methods into type bus so that when typist sees these messages coming off of Kafka it knows where to for these messages - great so how do we use this client well first we have to initialize it and then we say give me a cat named butters I get back a future and I map over it I get a butters cat and say you know I really like butters he's really fluffy and you know cuddly so I'm gonna go ahead and say I'm gonna adopt this cat I'm gonna announce it to the world by putting this message back on to Kafka all right so let's really think about how we can accomplish this request in response pattern over a Kafka so here I'm announcing to the world I want cats I put it onto Kafka the cat ran sees this request and we forward that the cat ranch will respond with a cat put it back onto Kafka and I'll get a cat back now if you ever used Kafka you're probably thinking yourself Rolo allo that's not how Kafka works and what is this bear doing on the screen and you're right Kafka doesn't work that way in reality so it's it looks more like this we have a me cluster that was represented by three different VMs and that request could start from me one so me one says I'm gonna request the cap I put it onto Kafka the cat ran sees that request sends back a cat through Kafka and then floors it to maybe me three because Kafka is managing the message distribution me three gets a cat on his lap doesn't know what to do with itself do I pet it do I feed it what do I do with this cat so in reality this is something that we're gonna be doing or what sorry something that we're doing to finish the response cycle is that we're starting up a a caca cluster actor on me one so me one is actually an aqua cluster so when we start up this actor it has an address and we insert that address into the request and that request goes to the cat ranch and the cat ranch Lords back the original requestors address so that when me three gets that message and that response it's it knows oh wait a second this cat isn't for me this cat is actually for me one so I'm gonna head and forward that problem I mean cat to be one and that's how we accomplish the request and response pattern and this is somewhat similar to how akka implements their ask pattern alright great so let's talk about code first data modeling with thrift Avro protobuf you first have to define a schema and then from that schema you're gonna get a pure or sorry clean Oh scholar object well that means that our engineers need to learn the semantics of this IDL and then we didn't want to jump that hurdle we wanted to move quickly so instead we want to work with just Scala objects so let's go back to the cat client definition fetch cat and cat are actually just case classes so how are we messing how are we sending these things over Kafka we're using a library called a verb for s and basically what this does is it will take a keys class serialize it down to Abril bytes type us will then wrap that in an envelope with metadata and attach a the canonical name of that clock in case class then serialize it again with Avro and for us because it's Turtles all the way down and then we will put it into Kafka when we pull those bytes off of Kafka we'll do the reverse well DCL has it back into the envelope and then we'll deserialize it back into the case class so that our client our code our business code only deals with Scala objects and doesn't have to deal with this DC relation you serialization dance you also notice that the topic name is the same as a canonical class name of the message that gets inserted to it and we're using the same Java convention naming convention to map it to a topic and this really allows us to manage our topics better so that we're not wondering what message is in this topic we already know by the name of it great so let's now let's talk about service discovery so service discovery basically might you might think okay well I need a new service so that means I probably need to spin up a load balancer I probably need to spin up some servers and I probably need to create a DNS name that other services can talk to you and then I need to reconfigure the clients to say like hey this is new service that popped up here's how you talk to it here's the DNS name here's the poor etc etc and that's a lot of hurdles when you're trying to move fast so what we're doing is we were relying on Kafka's message management to to remove that so what does this look like in practice say we are in production and I'm requesting cats and I'm getting cats I'm happy life is good right the cat ranch has provided me with those cats but then product kind of comes to us and says well you know these cats are pretty miserable they're cooped up in cages maybe they're too close to each other they're not really happy cats so let's come up with a free-range cat ranch and this free-range cat ranch is gonna have fields and it's gonna have trees is gonna have all the catnip that the cats can take and then they're just gonna do cat things right there's gonna be happy cats so we work hard we spin up this free-range cat service and Kafka manages the partitions so it now starts sending some of those requests to the free-range cat ranch service so what this means is that now the team can tear down the old cat ranch and all of a sudden I'm still getting cats I'm still requesting cats I'm still getting cats back and I'm none the wiser of where the cats are coming from whether it's a ranch or the free-range cat ranch and this this flexibility allows us to spin up and tear down services relatively quickly the last thing we want to tackle is configurable throughput management so in the financial space especially when we're talking about credit card debt at the end of the month at some given time we want to start paying off our customers credit card debt so we kick off all these jobs and when you think about these jobs they create a lot of data based queries what we don't want to do is we don't want to DDoS our databases and prevent our mobile clients from accessing the data because that might seem that we're down so we want to really control the throughput that's coming into our databases and the way that we're doing that is by using acha streams so you can imagine one of these jobs kicking off putting a hundred and fifty operations per second our database on the far right can only handle about ten or ten requests per second this is a good travel example that's a sad little database and can only handle temporary quests per second so we're using aqua streams to really control that rate and that throughput so that our databases has enough resources to also serve our mobile clients requests so that we don't seem like we're down or slow when these jobs are running and we can still satisfy the needs of our jobs but we can also still satisfy the needs of our mobile clients aqua streams is really great for also describing these complex graph or it has a excuse me it has a expressive grasp DSL that allows us to add some retry logic into the graph so what does this mean if for whatever reason a message is dropped because of network magic rule do whatever we want to be able to retry that message because that message represents someone's credit card day we don't want to miss a payment so we take that message we added back to a retry source and akka streams allows us to express a flow which will take some messages from the Akasaka source and some messages from the retry source and still maintain that 10 per second 10 operations per second throughput to the message handler the message counter is oblivious of whether this message has been retry previously or is a brand-new message all it really cares about is handling that message and making sure that operation succeeds great so let's let's go back and talk about all the things that we talked about so the first thing we want to do is have a convenient service updates and we do that by kafka Kafka is managing our messages and making sure that any new service can start subscribing to updates we have an easy and convenient API that can publish and subscribe and also send out requests and get back responses we have a code first data modelling with a Perl 4s and Kafka is managing our service discovery because we don't want to jump the hurdles of setting up a new network infrastructure whenever we spin up a new service we also have configurable throughput management with akka streams to really allow us to control the throughput of these asynchronous jobs all right so what's next we want to start thinking about distributed tracing with typos you can imagine a message being consumed by six different consumers and all of a sudden you're looking at a trace with one message with six sub graphs and those graphs can be graphs of graphs or graphs of grausens turtles all the way down so we want to think come up with a way to manage that distributed tracing and an inventing system we also want to start thinking about how we could use and leverage production data to load tests our staging environments you can think about maybe we want to replace our sad little database that's doing ten operations per second for something that's more scalable and instead of creating the data ourselves because you know we don't predict the future we can't predict how clients are actually using our applications we want to use real production data so that we're not missing out and in edge cases and we want to make sure that our systems that we deploy to production are predictable in a way that makes sense to our business all right I so that's the future of type bus if you have any questions these are the ways of means of contacting us I want to thank you all for listening to us thank you [Music] oh we have time for questions great five minutes for questions oh that was great dog thank you I had two questions one was for your idea of definitions you mentioned that you use Scala objects and then you serialize it using ever forgit whatever that framework was called but how do you share objects across different services that's a good question so right now we're in the transition of we have a mono repo and right now we are using SBT to build and compile that code we've slowly discovered that that isn't going to be scalable and so right now we're in the middle transitioning to basil but the way that we share that that code across our micro services is by using a mono repo and I'm curious are there any other frameworks in the industry that are comparable to what you just talked about to be honest I'm more familiar with Kafka than I am with Scala and octa and from my experience there isn't something out there like this that can fulfill a request in response pattern with akka are sees me Kafka they're so similar in the same idea other you do a refactoring let's say for example you're a fetch cat message how do you how would you rename something or in this case it's a really simple example but how would you add a feel or remove a field since it's gone into Kafka and the topic name matches the class name right so to make sure I understand the question correctly it's how do you handle the migration of the message itself right great so aro for s or excuse excuse me Averell is a version schema it has versions of schemas so as you add a field you can say that this is the schema that produced that field and it knows how to read it and create it back scholar case class and vice versa if there's a new field you have to make it optional to say if I'm reading it with an old schema just ignore that field and it works somewhat similar to protobuf but protobuf actually uses numbers where a bravura uses ordering but generally we're relying on Averill's we handle the schema migrations oh yeah go ahead yeah I'm not overly familiar with Kafka but we call correctly not if you're subscribing to a topic there are no like particular restrictions on who can read the messages and so I'm curious do you have any data that you're transferring over these over this bus that you only want certain services at your company to read and how do you prevent like other services from not accessing that data what form accessing that data gotcha so the question again is we have a topic and I want to make sure that only certain services are listening to that topic so there are multiple ways the easiest answer is just making sure that your code reviews are on point but the better answer is Kafka itself has role-based security so you can limit the amount of services that are actually talking to those topics by assigning those permissions to those services hi there I was curious how you were doing the load bound load testing and staging with production data and how you ensured you know no PII is being leaked and things like that great question so this is a thing that's in my mind right now and that's something that's in production it's a future project if you were to ask me off-the-cuff just off the top of my head how I imagined we would do this is that we would create a streaming app streaming kafka app that would potentially scrape out any PII data but also to be clear we don't quit PII data and Kaufman in the first place so once we have this streaming application we can then send it to a different cluster and from that cluster we can do whatever we want with it does that answer your question cool thank you there's one more right here yeah I was wondering if you can share some data on latency yeah that's a great question so again like the the app right now is in private beta but the latency that we are looking at depending on the base the use case is around 50 to 150 milliseconds and he depends on really on the workload and what the message handler is doing we honestly haven't spent a lot of time on tuning the cough code producers or the Kafka consumers so I think there's a lot of wiggle room in those perspectives but we're really not we're relying on our data infrastructure sorry our infrastructure team or what we call dev productivity tools dev productivity team it's really fine-tune and make sure that the latencies are acceptable to product so right now it's too early for me to say that you know it's gonna stay in those balance especially because the amount of data that we see is not really where we want to be so maybe we'll give another talk and you know the performance issues that we might see in the future um do you guys have any view to know where you request this in this topics how would you know sorry can't speak more into my how would you know where is your ear request is in all these topics that you have like like let's say a customer request comes in like how do you debug it how do you troubleshoot where the request is is it drop down in the middle or is it where at this yes let me make sure I got that question correct you're thinking about how do I trace back where this request came from it's that correct so right now what we're doing is every time a producer produces a message we're creating a UID with that message so we can trace back where that UID came from and what service they came from right now our app is relatively small so we don't have a problem of like thousands of micro services talking to each other I think in the future it could be more sensible to add like a service name to that new UID so that's more easily trackable but also we want to start investing and distributed tracing which will hopefully answer that some of those questions of where this request came from does that answer your question Oh gotcha so you're talking about like if the knesset handler is latent do you want to know where in the coffin topic yeah so I don't actually have a good good way I don't think there's a good way of keeping track of an individual message but we are collecting metrics on the end when we consume the message off of cough coats you see when was the message when was the message created and when's the time that I saw it create that diff and then put it on a graph so we have alerts on making sure that our consumers are not far behind or do we have time for more because I think you might have another one no okay sorry you know we're going to be right outside these doors let's talk more thank you all [Applause] [Music]