Akka - Enterprise Agentic AI with Tyler Jewel
Recording: Akka - Enterprise Agentic AI with Tyler Jewel
Hello everybody. Uh I'm Alexi Kraov, the founder and organizer of AI by the Bay, the longest running deepest technical independent open source conference which is being held for 11th year in the bay. It used to be called scale by the bay and data by the bay. We're doing different editions and today we have a very special guest tur who's the CEO of a company with a long history long relationship with scale by the bay and scholar by the bay. uh used to be uh called type safe and lightband went through an evolution and uh at its core it has very robust distributed systems framework called docker so now it's actually called one of his products so very happy to have tal here welcome and tell us a little bit about yourself and how what was your path with type safe light band and aka and you know why did you join a CEO now from an investor and how do you see the future of aentki with aa Well, thank you for having me. It's a real pleasure to be here. Uh, yeah, I I am the CEO of AKA and and the company's 15 years old. It started off as Typesafe, founded by Yonas Bonet, who is the founder and CTO and it was originally intended to be a development framework to make it simple to build massively concurrent systems on multi-core compute
And and 15 years ago, that was a really difficult problem. blocks threading. You had to have practically a graduate's degree in order to understand it. And and Yonis used 1970s computer science actors uh in order to implement his form of compute and and actors offer a really nice form of supervision uh and message passing asynchronous message passing in a secure isolated way uh which which offers forms of resilience and lightweight compute. Uh it turns out that AA uh as a multi-core compute unit works equally well if you think of it as supporting microservices that are operating over a network. uh and it became a full-fledged distributed systems framework where you actually push logic and data into the microservices and those microservices run in large clusters close out out to the edge and that offered some incredible uh outcomes for people who made use of the framework uh where you could get ultra low latencies for accessing your data either from reading or writing. uh 69s of availability because the system was spread all over the all over the place along with massively concurrent numbers of you end users because you could make as many replicas of the system as you wanted and uh I joined as an investor in 2019 a board member and you know really really saw an amazing potential with the technology in the company and and what had happened is we had a huge uh base of bluechip accounts there's about a 100,000 systems that have been built with A over the years one and a half billion downloads and on any given day two billion people are use an application that has a under the covers and it always seemed to just work but it was also plagued with some complexity and some challenges. So there was a long learning curve with it and working with Yonas and the team we saw a tremendous opportunity to transition it from a development framework to a platform uh a platform that could really simplify how you build these systems but also have high velocity with a very quick rapid delivery cycles and and that opportunity was just so big and massive that I I wanted to get involved with it and I I joined the CEO a couple years ago
Interestingly uh when we originally delivered this platform it was uh it was intended for application modernization that was the sort of use case. Uh but along came Agentic AI uh and about a year ago we had a number of our larger customers come to us and tell us that they had started to build a systems using Naka and we were very curious by that and and we became a little bit of some uber geeks in and around that particular domain and what we discovered was that Agentic AI is a distributed system. what you're dealing with are lots of agents that have to be coordinated over a network and and the agents have some sort of shared goals which is distributed state. So now you have a distributed system at the heart now backed by a stochastic LLM which is very unreliable. It was a perfect fit for us and and we've since pivoted and and kind of really zeroed in our platform to be specifically an agentic platform now. >> Fantastic. This is no this is thank you for this overview and I just you know might add that you know I've you know seen uh type A flatbed and aa folks going through this whole arc right because the server systems are very complex as you mentioned and for a lot of folks who come into genai right like they see tip of the iceberg they see few scripts in python and they kind of hope it's going to work if they unleash a bunch of agents right uh but you know to somebody who's been around the block a few times it seems to me like it's mostly distributed systems right so the the logic of interacting with the lens you send strings to it and it gets strings back right which is also you know better done in Java probably because it's more reliable but like if you look at this mix so you guys and I talked to Jonas right at J focus in February right uh I went to Stockholm as a community architect at new forj another Swedish company and you know s journals in a we had a conversation with him we had a interview right and we talked about this agentic vision. So I think it's kind of it's clear to folks who've been in the field that this is a match made in heaven but kind of you know there are some new folks how do you see kind of and what's the mix of kind of AI versus distributed systems? How much in this casual application do you see people actually spending time building distributed systems and all this logic? Well, if if you if you were to listen to the media, the the media would have you believe that everybody's building agents
And if they're building those agents, they're probably doing it in Python today. And and and actually, it's it's really not quite that is kind of a bit of a far stretch on the truth of what's really happening on the ground. uh what has happened is that uh desktop tools or desktop applications where you know you can build local agents that make use of local tools has you know really grown and and exploded. Uh but that's not the same as an agentic system that's going to run inside the enterprise. Within the enterprise itself, what we've laid witness to all is every organization on the planet either researching, learning, educating, maybe testing and prototyping, but a very very small number of them actually building systems that are truly agentic systems at the end of the day. Uh so it feels like that the market is still in its first ending in its earliest stages that's there. And the development that does happen, it's often times happened as a graduation of the data science teams. So you have these MLOps teams or the data science teams that were responsible for working with the data and the models
And of course, you know, Python is a great language for building models, training models and working with them. And so it was just like this natural extension where the those teams, those data science teams said, oh, you know, agents are going to call the models, so therefore we must be responsible for the models as well. And so you see all these Python frameworks that have emerged. Uh but what's happening in reality is that as these agents transition from prototypes to systems that need 24/7 operations that are resilient that never fail uh suddenly the CIO is raising their hand and saying hey wait a minute this is a different type of domain uh you know my organization is well equipped to handle that and as that transition uh takes place you're going to see the growth of Java and other languages that have kind of proven operation properties that exist for agentic systems. >> That's great. So, you know, the topic main theme of by the bay 2025 is reliable AI, right? So, we basically want to emphasize various technologies such as distributed systems. So, I remember I learned you know to huge surprise that era right like era is the most widespread Wi-Fi mesh and home mesh. I think it was one of the biggest AKA customers behind every era device there is an AKA actor right so I'm just wondering like I think you have a huge advantage over a whole bunch of this agentic new you know pythonic frameworks because they're aspirational like they they I hope you you're going to adopt them but Aki is already in production so it it seems much easier to add AI to already existed robust distributed system so I wonder if you can talk a little bit about you know some customers of Makov which already are out in the field and like like if they're doing PC's thinking of PC's like what would make good uh experiments for the existing customers
>> Yeah. Yeah. So we have about 45 customers in production with Aentic systems now which which is you know we're kind of thrilled. We're over the moon with that because these are these are hardened systems that's there and on top of that we just we just signed uh last week our first big win competitive win against Langchain. uh so in in an account that had only Python developers inside of it and we were able to convince them that you know a Java framework was the way to go. Uh so uh so our message is definitely resonating. Uh we're we're pretty pleased in terms of some of the accounts that I can speak to. One is Swiggy out of India
uh they do they built a personalization and a traffic routing system that uses models uh to do recommendations on that and and they had a subund millisecond roundtrip latency requirement. So they needed to build a custom inferencing layer. Uh so they used basically a giant aa cluster as as that inferencing layer. And another company who I don't know if I can mention yeah uh Tuby which is you know owned by Fox. Fox they built a real time personalization engine that will make content and advertising recommendations for users with this sort of continual feedback loop and they implemented agents that have over a half dozen models and it's sort of this sort of continuous real-time streaming mechanism that's also making use of models behind the scenes. And so that we're really proud of that. Later today, I'm signing a a new platform deal, new AI platform deal for the UK's largest payments processor. Uh so there, you know, so this is a fintech
This is probably the first fintech that is going to go down this agentic path. And and I think in the next um in the next two months, we're going to sign another almost 10 deals as well. Uh and these are all large enterprises, retail, let's see, uh who else? Uh health care, health insurance as well, u life insurance and uh medical devices. So these are all enterprises that are going down this path and most of them are Python shops that are now converting over to Java. >> Oh, I have a I have a personal question. I spent last three years trying to convince Python developers that they should not use Java frameworks as a software engineer. I uh was a Java developer, scholar developer. I used a actually in production uh back when I was working for Xedia
We used the actor system. But uh here an important thing I think to mention for our listeners is that you're not convincing Python developers to write Java bare hands right you you talked about the platformizing making a platform of vodka and uh that's that's what helps them to choose between length chain and aka is that correct? Well, when we when we personally sell into the enterprise, we have we have kind of a three-pronged value proposition that we give to them. And the and the pro the value proposition is we can get you into production quickly. We'll keep you there safely, and then we scale cost effectively. And and we can demonstraably prove that we are higher velocity than uh Python based frameworks, that it's safer, uh that it's more secure, and that the the cost of scaling is way cheaper. And so, let me just break those down and how we do that. And so the language the language that you use is irrelevant because these other values to the enterprise matter more at the end of the day. Now um how do we get you there quickly? The first is is that all the Python frameworks are just that
They're frameworks. And the thing about frameworks is that they still force you into composition of the system itself. So you have to piece together the agents, the memory, the orchestration, uh if you want streaming or APIs. These are all different piece parts and you got to find a way to make them work together. They have different development approaches, testing philosophies, um, integration, CI/CD mechanisms, packaging and deployment. With AA, we take a systems approach now and that we look at the entire, uh, Agentic system as a uniform bundle of a service that you're going to deliver against. And our SDK gives you declarative ways of defining what that system is with very basic Java programming. But more importantly you we uniformly approach development unit testing evaluation integration testing packaging deployment all as a single unit
And when you treat it as a high velocity unit that way uh you end up having a lot fewer bugs. Uh the the Java ecosystem for modules and third party dependencies is safer. Uh it's it's a more trusted uh bomb sbomb that you're able to build. And then because there's also a standard way of packaging it, uh you have mechanisms where you can do rolling updates with this in real time as well. And so so this is all this all this benefit comes down to the fact that you can make lots of improvements on a very short order. Now we'd actually argue that our SDK syntax is as simple as the Python syntax now and and because of the way that we structured our SDK, it's so tightly constrained. There's only one way to do everything in our SDK, which means that we can feed it into AI like Gemini or claude code or whatever it is. And the AI now generates entire systems without hallucinating
So you can get started with this and start building. So that's the get it there quickly. Uh I can elaborate on all the things we do to keep you there safely, but it's about having a compliance and regulatory process that allows you to build trust. There there's a very special technique in which how you trust these systems. uh the systems are based on unreliable networks and unreliable LLMs. So trust and governance is so critical on this. And then on the scaling more cost- effectively because of the JVM and the way that the bite code interpretation works and our actor model, one co core of aa is five to seven cores of python um in lang chain. And so when you're looking at your cloud compute costs, um you're saving anywhere from, you know, $2 to $4,000 a year just on the raw compute cost because of the efficiency and the density of the of the systems that we can run
You put that in front of uh hiring managers, executives, and and it becomes a a choice between do we want to go with the language that is the most familiar or do we want to go with the system that's going to be the cheapest and the safest. So from my experience a lot of times people are not like Python developers are not afraid of the language itself but more about the GDM uh ecosystem things like garbage collection and uh overall uh distribution etc. So this is this is what I think they struggle more with because this the syntax is easily you can easily fight that right now with all kinds of copilots versus like with a distributed system. Yeah, >> to a degree. But with with with AA, we've hidden almost all those complexities. You know, garbage collection, threading, asynchrony, you know, all all those things. And what ends up happening with the Python developers is is they they think that it's really simple to get started, but the moment they want to have any sort of concurrency with that, they're they're also now wrestling with the global interpreter lock. They if you want to run multiple processes concurrently, it's like a megabyte of overhead for each one, whereas it's like four kilobytes of overhead for actors
And the Python runtimes aren't inherently resilient, meaning that if there's failures in processes, there's got to be some other supervision that is managing against that. And and those are all just solved problems um inside of the frameworks that we have inside of the JVM today. So really this SDK, you know, we've simplified it down so that yes, you might have to learn a new language, but now you're just a bunch of a bunch of day two issues that you just don't have to worry about anymore. And that and that wins when we can get to the bake off. the bake off plays very well. >> That was a little derail. Sorry about that. >> Yeah, it's all right
>> No, [snorts] that was a good good question. I I have a couple followup questions. So, first Tyler, again, like this is the first conversation we have and I knew you as an investor from Dell Capital. Like you're very deeply technical like can you tell us a little bit why did you invest in light band and like what is interesting to you in this space? How did you come to play in the distributed systems playground? So you know all the issues and costs and technology and so forth. >> Yeah. I mean I I ultimately invested originally as it was 2019 when I invested and there was two reasons that at at the end of the day which was one is there was just no other technology on the market that could make the guarantees that aa aa could make and and it was proven you know not only could they make the guarantees but it was proven. I mean then you know everything from Starbucks to data dog Walmart Apple you know you you name it. we've we've kind of run some of the world's largest distributed systems at this point
And then the second reason was that there seems to be like massive adoption at least at the time. It it felt like there was a lot of massive adoption. This is still when the product was still pure open source uh and a lot of that adoption was coming from the scholar community. And so the the investment reasoning was unbeatable technology and and probably a really big market. at the time the business model didn't work out the way I thought when I first invested into it and that the adoption was there uh the technology was definitely proven but because it was open source it was very difficult to convert all the adoption into a commercial business and so that that didn't play out against my hypothesis the way that it did. Uh and we've we've had to basically to solve that problem, we had to move away from being a development framework into a pure platform platform plan. Uh stop selling to developers and instead sell to uh senior executives that have to pay buy pay into outcomes and then we ultimately change the licensing model from open source to non-open source as well >> which follows majority of SAS offerings right so this is not surprising this is like pretty much a norm. So another thing I I had a question uh I think you mentioned a fintech P and that's interesting to me because I noticed that banks are very conservative they basically run on GVM right and so it's interesting to me that you know for compliance reasons security reasons right they're very like you cannot just bring a jar into a bank right like it will be put on the surgical table disassembled looked at put together right like it's not as easy as people think you cannot people install stuff right you cannot you know use maven even to install stuff you have to go through you know jump through hoops so to me it signifies that you know it may be actually very good pre-existing condition that fintech is run on JVM so I knew you know companies like Tali you know they actually selected JVM intentionally right and others to like if you are in a startup in you want to be in JVM because you're going to eventually exit by selling to a big bank which already runs on JVM so do you see this happening and how do you think is going to affect adoption of agentic frameworks? >> So, uh we're we operate in 52 banks around the world and about 28 I think it's 28 or 29 of those use us for payment uh >> on that
So, so a large number of the world's payments are AKA influenced at the end of the day and what really defines uh you know the fintech environment is really a fear of loss of data as you say and so they absolutely need guarantees that um that there's never going to be any sort of data corruption and on top of that meaning you can always recover from any possible anomaly that's there and then you've got to layer in the fact that no outsider or no one who is not authorized to have access to that data should ever have access to that data. And so those are, you know, those are really the things that drive them at the end of the day. Um, and the JVM and it's not just the JVM, but it's also the way the JVM ecosystem when you look at how all the different vendors go into um contributing to the JVM and then all the modules that are on top of that, they are uh it's kind of a trusted ecosystem that has been managed by the commercial vendors themselves. It's not by uh open- source developers in that sense. And so the supply chain, if you will, is as safe as the um as the application inside of it. And so it's a combination of like frameworks like AKA that have been able to prove and demonstrate that it's impossible to corrupt the data that it's impossible for us to get access to that data combined with this sort of safe secure supply chain of all the modules that make it possible for the FinTech to make use of this. I think that the fintexs are a recognition of what's going to happen across all the enterprises which is uh I I think that at least half if not threequarters of all the of all the agentic systems will probably be built with Java. I think that the models and the model training and the model inference will probably continue to be Python because there's no need for that to be a JVM
But the applications that make use of those models, there's really no reason why you would want that to be Python. You would almost always want that to be a JVM. >> Absolutely. I mean most of this you know training setups are commodity and once you have the model the real question is what are you going to do with it right and like there are many ways to run models and there are many providers competing on cost so it will effectively become a commoditized resource again is a fellow Swedish employed person so you know aka is a name for folks who don't know it's that's the mountains in in Sweden right so the original team named it and I'm now you know forj which is another Swedish company also on JVM. J stands for JVM for folks who remember this. And interestingly enough, right, we have a very similar situation where there is, you know, new was the original graph database and now there is a million graph databases, but new forj dominates in the market of very one simple guarantee that we're not going to lose your data, which is kind of important, right? If you are a database, you you can be very fast, you can be very easy, you can be very convenient. In the end of the day, if you never lose data, like we'll compromise on speed, we'll compromise on whatever, but we'll lo not lose your data. Another interesting parallel to me is that the enterprise version, we have an open source community edition, but the enterprise version has all the permissions features like you mentioned, right? Like which is really important
So I can see a lot of so maybe you can talk a little bit about the origins of A and Yonas Bonire, right? who is the visionary behind aa I just wonder like he's the CTO he produces he spoke at scale by the bay you know he's a creator of reactor reactive manifesto can you talk a little bit about the kind of intellectual you know capital behind aa all these kind of uh things which others did which led to this point in time I think what uh is the brilliant part of of a's principles is that if if you want to build a system that never fails you you really have one of two choices You have you have one or two basic design philosophies you can take. One design philosophy is to maintain the state of the system as a durable record that is always accessible. Right? That's the temporal approach, the durable execution approach, right? Because you can maintain the entirety state of the system um as a durable record, you can always go back and recover the system and then continue on. So, it's provably works, but the downside of that is is it's really slow because every every kind of state update requires you to go to some central central point and go get access to that and and then it becomes a single point of failure. Uh so simple to understand but really slow at the end of the day. The other design philosophy is to say no, the application itself needs to take responsibility for its own outcome which means that it needs to divorce itself or decouple itself from all of its underlying infrastructure and has no dependency and so that it looks at persistence and compute as uh commodity resources that are available to it. uh but and and this is the approach that Yonas take and so when you take this approach what you start doing is I'm going to design a distributed system where you make the assumption that that uncertainty is abound that whatever can go wrong will go wrong that I cannot anticipate whatever is going to happen and that the system itself has to be independently capable of uh observing the problem and then adapting uh to whatever that problem is presenting itself whether that's a hardware failure or a network disruption. And so when you do that in his, you know, in his logical mind, uh, you have these distributed nodes
They're running in these large clusters, but the application itself takes responsibility for all of its stateful data. It's not just responsible for its compute. It is the system of record for that data. So your data actually runs in the application. Uh, and then a will persist that data behind the scenes. But that's a responsibility to manage that. where that data goes, how that data goes. That's an independent concern
And so the application itself becomes its own living, its own living organism. It has responsibility for its own SLA and is constantly adapting towards achieving that SLA. And when you do that, you get a system that's effectively self-healing. It can move from one location to another because it might need to be able to do that. Um, and then it also knows how to scale itself out. This becomes very elastic back and forth. And so what you end up getting is a system that can hold responsiveness um indefinitely. And when you can hold responsiveness indefinitely, you start unlocking these amazing these amazing SLAs's and these outcomes for that
And that's and that was Yonas's approach and that and that worked. >> Fantastic. Thank you for this overview. Uh and I guess the last question we should have right uh to run it up is uh obviously things are moving very very quickly. I think you guys like enter the field hit the ground running built on distributed systems kind of adding aentic AI components others are basically starting with the other end where do you see this field going and as practitioners and as attendees of AI by the base are learning conference where should they focus to kind of build the agent future which will be the one which we're going to build against well I mean I think that where most people are at right now is they just need to keep doing experimentation and and understanding what the different frameworks are that are out there and what their different direction and goals are. But you know what's going to happen here is over the next over the next 5 years I think we're going to see millions of agents that get developed and those agents are going to probably be running in islands pockets pockets of intelligence that are throughout the enterprise and they may be in different frameworks they may be under different jurisdiction u whatever that regulatory regime might be and in order for an organization to truly unlock the power of agents you're going to be needing to find a way to combine those agents into some sort of cooperative ative mesh where you know your your employees can come along and say I have this problem I need this goal I have this goal and I'd like to be able to build an agentic system uh that makes reuse of the agents that are already there but towards this sort of higher level goal this higher level of abstraction uh and so that is a type of orchestr enterprise orchestration that is a planning system that works towards goals that works with all your islands of int intelligence that's out there. And so I think that the future is this enterprise cooperation system. And really in order to deliver against that, there's going to need to be some sort of agentic uh control plane, a global agentic control plane that can discover discover these islands, interact with them, uh combine them together and and enable this sort of kind of goal, these goals and outcomes that uh that the enterprise wants to achieve
>> Does I think that's actually sets a very good action item for our community and we can talk about it at the conference, right? How do we interconnect all of these islands of intelligence? And probably that's going to be a distributed system. So, and you guys have a lot of experience with that. So, thank you very much. Looking forward to your keynote and uh participation in AI by the Bay. Thank her.