Devreal

Rust AI & Data Meetup Introduction: Alexy Khrabrov

Event: Rust AI Begins!

Rust AI & Data Meetup Introduction: Alexy Khrabrov

Recording: Rust AI & Data Meetup Introduction: Alexy Khrabrov

Welcome to AWS Builder Loft. My name is Abrao and I'm a founder and organizer of this new series called Rust AI with my co-organizer Scott and Manvika. Please show yourselves. Uh and uh this is a super interesting moment because I've actually been in three incarnations of this Loft. I run Bay Area AI, which is the oldest uh living in-person AI meetup in the world, which is operating for 15 years. And also conferences like Scala by the Bay, Scale by the Bay, and recently AI by the Bay. Who was at some of the meetups by the Bay? So, we have some folks. I want to specifically give a shout-out to Jason Mars, who was the co-organizer of the Scala by the Bay in 2014 and uh various meetups in that community

So, you will see the kind of the parallel why I'm mentioning this, right? Uh so, I just wanted to kind of give a very quick kind of opening uh like a history flashback, right? Which will explain uh why we do this meetup, how we want to do this, kind of what is the tradition we want to base on, right? Um and because those who don't know history are condemned to repeat it, as we well know. So, uh but first of all, I want to thank uh there's a lot of folks who worked hard to make this happen. Uh first of all, uh AWS is, you know, the this location is one of the best in the city. There is not many companies who continuously open offices to meetups after COVID. So, thank you, AWS, uh for making this happen. And uh Manvika is at AWS Loft Key. She's going to give a talk about it. And uh there is a lot of folks actually I met, you know, who run this Loft and it's a lot of work

It's busy every single day, right? And like continuously in operation. So, uh huge contribution. Data Bricks sponsored food and AWS sponsored drinks. So, thank you for that. If you go to the machine, you can pick any drink and you will just check out for $0, right? So, that's That's very nice. It's really good feeling, you know, very unusual in America. So, I like it. Right? Just try it

You will have a kind of a kick out of it. Something for free. Now, and so I just wanted to give like a very quick kind of background how this came together. First of all, I wear many hats. I am the head of ecosystems at company called Lake Sale and our CEO Shai Harel is going to give a talk after me. And basically, you know, one day they just woke up and decided to write everything in Rust and they rewrote it in the whole Spark ecosystem. And I was actually very interested in Rust myself for a long time. So, it gave me kind of a forcing function to do a lot of open source engineering, which I kind of postponed

And now with AI is possible. So, everybody who used to be a software engineer, then become a manager or a director or somebody else, it turns out it's actually very good time to be an engineer. And so, this meetup is going to be an engineering for AI. So, engineering is very much overlooked. A lot of talk we have in this town are about agents and they're not about engineering. So, I run AI Agent itself, which is the biggest agentic AI meetup with colleagues from Formal IBM and Neo4j and Cogney and Claw Max. So, so we do a monthly meetup, right? And what often happens, we talk about all kind of stuff which is not really engineering, right? And it's almost like, you know, markdown prompts are going somewhere, something happens. And it's kind of it's very difficult if you're an engineer, if nobody explained from the first principles how is it exactly built? What does your observability mean? Because all these a lot of the startups install a blob which inside of it has 50 different things including Kafka and other stuff, right? They hide it from you, but you cannot do observability without distributed systems without full understanding how distributed system actually work

There is no way around it, right? So people who try to hide it and oversimplify it basically set you up for failure later because performance considerations you're not going to hit them and a lot of them do not have any customers. So their observability shows nice graphs because they don't actually hit any loads. So I think it's very important, right? To kind of go back from first principles. And so if for me an example here is company called Pydantic. Who knows Pydantic? Who uses Pydantic? They're not here in the room which is a pity, but we'll be they'll be here. So Pydantic basically put a Rust-based type system inside of Python. When you're using Pydantic models, you're using Rust. You don't even know this, right? And they use Rust to implement a bunch of things behind the scenes

So they have a observability system called Logfire. Uh and a lot of people are using it. So I think this is kind of a proof point, right? That sound engineering is actually required for AI. And if you go to a Pydantic event, they have the slogan AI is still engineering. So I'm basically going to borrow it, right? So here's a little bit of our about myself. So So I have a PhD in computer science from UPenn and my advisor was George Cybenko. And actually did a lot of work at Dartmouth. And so I did a few things in AI before they became fashionable

So 1999 we put out the foundational paper called Intelligent Agents. And that was a blueprint for agents. So George Cybenko is a mathematician. He proved so-called approximation theorem in 1989 that any any function can be modeled by neural networks. So he's a kind of theoretical grandfather of deep learning. And so he he had this very good mathematical view of the world. He He created some of the earliest quantum computing theorems. Also, he he's probably he's ahead of his time by 10 years

Right? So, we did this in 1999. In 2007, I wrote a paper called A Language of Life, where I took the MIT reality data set, where they modeled how people are walking around Cambridge, and at that time they only had Nokia phones. So, they could only record your trajectory as a series of integers, which are the IDs of the cell phone towers. Right? And so, and I was very interested, can I detect your trajectory as a language? Right? So, when you walk along the river and you turn towards a coffee shop, and you sit down, and you walk like through a park, or somebody likes that, somebody is kind of romantic, and somebody doesn't care, they just walk through the buildings and offices, and they just blow ahead to the straight line. So, and I found that indeed, you can determine every single person from this data set individually by a hop of four towers. So, each of you has a unique segment. Probably around four cell towers, right? A few city blocks with a specific to you. So, I kind of, you know, asked this question

Again, I had to use a C++ package for n-gram modeling. It was before large language models. But, you know, the the kind of the questions were there. And so, then I wrote a thesis. My PhD thesis is called Mine Economy. And the question was, can you model influence in social networks? And so, I kind of bootstrap this economy where you exchange information with somebody else. Let's say Obama would sit you. Obama has a lot of influence, you have nothing

So, your influence actually jumps dramatically, right? Somebody important notices you. And Obama loses just a tiny fraction, right? Like it's really nothing. So, there's a proportional economic exchange. So, I bootstrap this economy on Twitter as a graph, and I discovered Justin Bieber. So, this is my small claim to fame. When nobody knew who Justin Bieber was, I actually knew who he is. Talk to me later if you want to know more. Right? And so, but here's we come to programming languages

So, I thought like a PhD, I already had my first son being born, like my second daughter on the way. This is the only time in your life when you can actually like really do whatever you want, all right? You're at leisure. So, I thought like I'm going to learn Scala, Clojure, Haskell, and OCaml, and I'm going to build my thesis not as a paper, but as a living system. You can still find it on GitHub in all of these four languages. Guess which language was the best? Which was the fastest, easiest, most performant for the task? Anybody? Just shout. Any. No. No Haskell

Haskell failed miserably. Anything else? Scala. OCaml. Exactly. OCaml is was my go-to thing because it's the speed of C. It does not like dogmatize in, you know, it's not like Haskell, it's not religion. It's hash table is mutable. So, the graph was modeled as a hash table, which is extremely performant and mutable data structure, very suitable

Haskell map is lazy. If you want to serialize a map on disk, you will wait forever. So, you have to string strictify the map. So, it didn't really work. So, that was a lesson, right? Like if you choose the right tools for the job, it dramatically simplifies your life. So, then I came to the SF and I was going to start doing Ruby, and they were Googlers, so they spent three days writing tests. They did TDD in two days writing Ruby. Eventually, they had 30,000 tests, and they never finished the tests

And so, I just wanted to Scala. So, this is how I came to Scala. I built the biggest Scala meetup in the world called the SF Scala. In 2012, I invited graduate student from Berkeley, Matei Zaharia, to talk about this project called Spark, a little project, and he presented that, and then I asked him to do the first Spark meetup. So, we were two co-organizers of the very first Spark meetup in the world. And I think that was the moment in time when Scala and Spark connected because everybody who wanted to big data could just naturally do it interactively at the prompt. And it just felt very organic. Like you say map, filter, reduce, it's basically just there

Right? So, that was very important to understand that the language enabled this whole big data revolution by answering what developers intuitively thought should be the operations that they want. So, that was I think the biggest Spark contribution and later at um Scala by the Bay conference, Martin Odersky, creator of uh Scala, had a talk that Spark is basically Scala collections distributed on the cluster. All right, so I created this conference is called Scale Data and AI. By the way, the last one was in November. We run for 13 years. It's the oldest uh independent conference in the area. So, um and now we basically arrive at this moment, right? So, basically my claim is that Rust is the same kind of uh but different next iteration moment for AI. Rust is the language which is uniquely positioned to power all the AI infrastructure because Rust is the foundation for Pythonically layer on top

If you look at what Python is doing, Python never does the heavy lifting. Python is calling into a blob of native uh code which is normally, before Rust was C++, so PyTorch is basically a giant C++ system and Python is calling into it. Previously, it was Lua. The original Torch was Lua, right? And uh not many people know Lua, so they replaced it with Python. Suddenly, they have the best deep learning framework. So, uh Rust is already doing a lot of the same things. You are calling into Rust. So, for instance, we have UV which runs Python build tools instead of pip

You have right, Astral is doing this at um uh Frontier Labs. So, uh and from these talks, you will see that a kind of uh evolution. So, one of the points of this uh first uh meet up is to have a kind of ground review. Let's let's have many talks representing multiple corners of the AI and data ecosystem, right? And kind of see where the landscape is. And the hope is that it will lead to this ecosystem network effects. Hopefully, you guys find each other's projects useful. If you have ideas, if you're building something, submit talks, you know, come and speak. We're going to All of these talks are short by necessity because it's an overview, but we hope that every single speaker will come back for a longer exposition

And all of the materials abstracts, slides will post on rise.ai which will be the blog. So, you know, as opposed to kind of transit Luma events, we'll have the permanent record. All of the talks will be recorded, posted on YouTube, and linked on the meetup. So, that's kind of the story, right? And with that, first of all, I want to again thank Scott and Money who are the co-organizers. They're around, right? And thank you very much.