Transcript: Ritchie Vink, Polars — Interview with Alexy
I'm Alexi Kraov, the organizer of Rust AI community. We just held our first meet up in Amsterdam before the PI data Amsterdam Italian office and we have with us Richie the founder of Polers. Yes. And you guys chose Rust for the implementation of Polers. So tell us why you selected Rust? How does it help to build the AI infra structure of tomorrow? Well, when I started a polar, it was mostly because I was interested in ROS. I was just very excited about the tech, about going into more low-level languages, uh, but also in a modern language. Uh, that was 6 years ago. Nowadays, it's also proven to be a very good language for AI. And, um, yeah, that's mostly a lucky coincidence, I would say. uh for a lot of people rusts were seen as the the complicated language the the hard language and now this whole strictness this whole borrow checker is actually very powerful at compile time you can prove uh a lot of invariance uh that your thread safe that your uh memor is safe um so you can yeah you can uh let an agent code with much more confidence so basically now people do not really have to fully understand what's going on because if you have a general idea for a good engineer and you can unleash your agent on Rust, yeah, it will be better than equivalent blob of Python. Yeah, in Rust I found it to be a lot better than it is in Python. The the agent has much more information uh given back from the compiler. I would still say you need to be a good engineer to get them the best out of it and getting in the loop is still very important in my opinion. But uh yeah it was a a lucky shot in that sense. So uh polars right? So basically like there is pandas there is polars there are data frames in spark. How did you come up with the idea of polars? What makes it special? Uh you have a growing community like what's what's special about polars which makes it a good choice for data frames? Well when I started polars you you only had pondos it's a def facto dataf frame library. Um and pandas was very uh powerful in a sense with a very rich API. You could do a lot of stuff but it was not very predictable in the API. Um everything was quite local and API could there was no consistent data model but that never made a lot of sense to me. But also pandas is in the business of of data processing. However, it didn't look as much to databases or query engines in general. It didn't make a query plan, didn't make didn't do any query optimization. Uh it utilized NumPy or Python itself for its data uh storage. Uh whereas there's a my there were sort of uh decades of of research uh within databases that was sort of ignored. Uh I think it was Dundas was mostly uh originated. It feels like it just came to be and whatever was out there was used but it wasn't uh connected to database research and when I started writing I wanted to be to be multi-threaded I wanted to be lazy and I wanted to have very optimizer um and the more I learned about it the more I felt like this should exist for a single node I want to get maximum performance out of my laptop because pend was kind of the design way to slice and dice tables right so it came from the need of data scientists not uh from the background of database engineers who properly can populate the tables. Yeah, but there's a in the vend diagram you need both. Also, a data scientist was waiting for it was pretty common to be waiting for a join in pandas and you you just saw one core in your your uh system monitor uh chugging away for 10 minutes. Why not use all the cores? Right. Right. Right. I'm very curious that you said about not predictable because I've been you know early organizer of scala communities and spark. So spark is effectively in my view and marski had a talk at my conference scale by way that basically spark is distributed scola. So scola collection was the thing which gave birth to this modern data processing where you treat everything as a collection and so the the winning uh of spark it you know had this moment because everything you would think about in the API map filter reduce you know it's object-oriented so something map would map something filter would filter so do you mean predictability in that sense like you think of the method it should exist on data frame and it should be there with a proper name yes but um because we make a logical plan. Um we had to come up with a data model and it should like relational model. The relational data is a logical uh if you if you have an API and you make a method on top of a data container, but that method can just do local implementations and doesn't go through a a a engine. It's very hard to make that consistently behaving. But if it is forced to go through the engine through a virtual machine then you have to come up with a data model. You have to come up with something consistent otherwise right it will not work. So you so you mean predictability in terms of consistent performance consistent performance but also consistent behavior consistent types. Yeah. Um what um pandas was non strict. So it used to be that if you uh parsed a an example it had a column of strings and it had to parse a date column out of that you had to parse that as a date column. What it used to do was run Python daytime parsing on every element sequentially. Uhhuh. But if the first one was date month years and the other was years uh month day. Yeah. It just swapped them around and you had or the first one was milliseconds, the second was seconds. You had not a un you did not have a unified data type on a column. Mhm. Uh you would have different data types within. So you didn't really type the column properly. No, it was not consistent in its type model because Python is not really a type language and so people don't think in terms of types. Yes. But it could have been in pandas. That's right. That's right. I mean Python is just the host language. Correct. You can make anything out. Correct. But because it comes from that background, people who did this did not really think of imposing a column type. Yeah. And um my experience was also that my types could change in production workloads depending on what kind of data went in there. Okay, I had the same query, but I had a rolling window of data. And all of a sudden, I was not indexing with a integer, but I was indexing with a float. And my production pipeline broke uh 40 minutes. I will be terrified as a strongly typed person. Yes, I will be terrified of this. So these things, yeah, these are paper cuts that hurt me and I wanted that uh fixed in a data frame library. Interesting. Interesting. So basically, did you start with Rust right away out of the gate? Yeah. Okay. Yeah. The initial goal, so Paul's story of just moving goalpost. The initial goal was, oh, I really like this this library or I really like this this language in Rust. It doesn't have a data frame library. Let's make one. Uh I worked my first join. Uh did a benchmark with pondas and my join was super slow. So the the next goal was make it faster than pondas. Mhm. Then the next goal was okay make a python API. Then initially I wanted to copy the pom API and then I learned what what's the consistence behavior? I could not find it. Mhm. Uh and then I learned about databases and about lazy programming and then at that point I learned about the optimizer and I thought okay this should exist. Mh. Uh but it was constantly moving goalpost. So the initial goal was just to across the library. Currently I want polars to be the fastest engine for any scale. I also wanted to be the fastest on distributed. Mhm. When I started this was never my goal. I would have laughed in your face if you said we would be doing this. So um yeah. you are basically growing with the community and we're here at the community conference. I'm super curious. It's PI data. It's Python. We're in the same boat at lake sale. We have you know rust engine but we uh most of the customers talk with through pi spark because support spark connect right. So this is kind of so we say pip install pi sale. So how do you interact with the python ecosystem? uh and like I think now a lot of folks you know in in Bay Area where you know I'm from basically solentic and so like they all agents are in Python. How how do you see uh polar kind of being in this Python ecosystem and specifically in the ecosystem which is all Python now? Yeah. Yeah. So Polar is a Python first library. We focus mostly on our Python API. It's a Python data frame API. Um so in that sense we're good. uh we're in the corpus of all the big the big models. So models know how to write polars. Uh a benefit is that at before we run the query you can also do we can already do a lot of type check. So we so you write a query you the agent gets feedback on the types or errors before it runs a query. So the wrong query will not go and reach the Yeah, there are still of course some runtime checks still but we try to catch as many errors as possible up front. Um so that's really beneficial compared to um uh compared to something like pandas. Um we're also focusing more on SQL now because agents took that friction away. Humans I mean humans don't like to write very complicated SQL strings on typed an agent doesn't care. So um yeah the world is changing so uh looking into that as well but uh I think there's still merit in in a API especially with regard to composability. You can pass expressions to functions let them generate stuff. So the meta programming you can do in polars is very powerful and it's all typed. Yeah, very cool. And so the last question over here at you know it's your hometown Amsterdam right where Polaro started this is one of my favorite conferences PI data I'm coming like year on year uh what's special about Amsterdam data community like what's kind of you know makes this conference you know good for you and page in general well it's a it's a home game of course so that's very easy I think Amsterdam is great because of all the the international And it's a very it's one of the largest pas I've been to. I've been to multiple padas across the globe. But um yeah, it's always a big a big It's a big one and one of my favorites too. Yeah. Yeah. And uh just a lot of uh I think different uh business areas, different different fields. So not all agents like they're actual businesses using data for something, right? which is and uh them uh in the Amsterdam is a great city to uh to have fun. So uh one of the best one of the best and is there like a lot of customers local Netherlands customers who you see bowlers? Uh definitely definitely a bit less than in the US. They our market adoption is biggest in the US. I think the US is just Europe is always a bit slower in adoption. they look at okay what's what's big in the US and then 10 years later it sort of moves over but we have a lot of photo users uh uh in Europe as well um nowadays luckily in the beginning it was always uh we were much more uh well known in US currently it's uh uh we're also big in Europe but yeah yeah it's spread in the world and you're one of the few remaining European companies doing Yeah. Yeah. Yeah. Um Yeah. So, um I think after the the DB acquisition, now we're uh yeah, we're the European tech company. Yeah. The independent Amsterdam different company. All right. Well, thank you so much. We're looking forward to hosting you at the Rust AI events in Bay Area. Cool. And I think like, you know, we'll have, you know, like a system of people doing basically what they call AI infra 2.0 Z where we power the fast, you know, AI of tomorrow with all this wonderful Rust engineering. Yeah. Yeah. It would be great to uh to be there and and be in the Bay Area. I always like to have this whole um vibrant culture. Looking forward. Cheers. Thank you. Bye, guys. All right.