Transcript: Matt Topol, Columnar, on ADBC — Interview with Alexy
Hello everybody. I'm Alexi Kraov, the organizer of Rust AI and head of ecosystems at Lake Sale, the company which wrote Spark and Rust. And here we are with us we have Matt toel, co-founder of Columnar, uh PMC of Apache iceberg and arrow and arrow and magpie recently right and Apache Macpie Macpie we're going to ask about that and so Matt is uh working on ADBC. He gave a talk on ADBC and DDB here at PI data Amsterdam where we are on location. So it's a wonderful industrial space. You hear some echoes of Python uh echoing through the building and uh uh super excited uh Matt told us uh cool things about ADBC uh kind of becoming the backbone of uh what I call AI infra 2.0. So tell us a little bit uh uh for folks who don't know what is ADBC, why is it cool, why are you going around the world? Easiest easiest way of pulling it out. I mean I'm trying to kill ODBC is the best way to describe it. If you're familiar at all with ODBC or JDBC, you know data connect, you know, how you connect to your database. Uh historically and traditionally, all of that connection is row oriented. uh because even though all of our almost all of your data analytical systems and analytical databases are column oriented. So you're paying this huge transpose copy cost of getting your data from your uh from your data your analytical system to uh uh uh uh to your local system to then your data frame libraries of your choice. You know polars, pandas and so on. And you have to transpose it all from columns to rows for the compute for ODBC JDBC. You transpose it all back into columns because of your data frame libraries, right? And so ADBC is a similar concept to ODBC and JDBC. It's a single interface client API and you have drivers that implement the API and the big difference there being that it the data goes through the interface using Apache Arrow which is an in-memory column oriented data representation format. Mh. So your data stays column oriented the whole way through end to end. So if your data system returns returns column oriented data, it gets to stay column oriented. If your data system is say Postgress or something that doesn't return column, it returns rows anyways. Well, your application still gets to have a consistent column oriented representation. You don't have to own the type mapping and all the logic and build all the logic for actually doing that conversion back. It just makes everything simpler while still maintaining a significant performance improvement. Yes. And you mentioned that uh it's not just like yeah building uh self-contained ADBC. People who install OBBC have a hard time finding configuring and doing it and you guys have a CLI to do that. So so so first of all ADBC itself the interface is part of the Apache Arrow project and therefore is completely open source Apache governed and so on so forth. Mhm. Most of our drivers are themselves also as good as what it is. Also, you know, open source. And so if you've ever had to figure out how to find, install, and configure an OBBC driver, you know, it's sucks. It's hell. It's awful. You go, you know, you go to some vendor site that you hope is the right one. You download a package that you hope is the right package. Place things in different place. Like, it's insane. Um, it's it's insane that that's what we've been doing for 30 years. So, so with columner, we built out a CLI we called DBC, which we like to say basically it's UV for drivers. Mhm. You know, it's a package manager to manage your ADBC drivers. Makes it super easy where you can just do DBC install BigQuery or DBC install, you know, flight SQL, right? And then you can connect to your favorite data system using ADBC because it just downloads the driver, installs it, and then you you can just, you know, DBC install flight SQL and then talk to sale or talk to Spice AI, you know, just using the driver and the driver manager. And so it makes everything just more seamless and a better developer experience. Mh. And also gives you more confidence in the supply chain because all the packages on the CDN are themselves signed by us validated. DBC is going to verify the signature when it when it downloads and installs it and so on so forth. Is it a binary install or is it compiled from source? Binary install bin install. So it knows basically for given it uses your it uses the it uses the system architecture and OS that you're running DBC on. Mhm. and grabs the appropriate uh binary for the platform you're running on. It's just we are currently we actually are currently someone started contributing to DBC a platform uh option so that if you really wanted to you could tell it to download the binary for a different platform than the one you're running it on if that's what you want to do say for doing a docker image build or something like that right it like cross compile kind of exactly it so cuz because we we when we when we deploy the drivers to the CDN we build the drivers for the different platforms and put it all put all the different platforms in the CDN and then you know just decide which one are you going to grab when you install it right right so uh it's interesting right so I think it's kind of the the contour so the AI info 2.0 zero is shaping up right it's fast it's native it's columner everything is columner right so kind of the uh it's it's interesting uh that we do not need to incur now the penalty of converting from rows to columns and back if we have columner for all so with a columner so we natively kind of align with this and also I just learned you told me recently that you guys have spark there is the spark adbc driver you can just do dbcin install spark y and when you're building out the u the URI that you tell it you how you connect. You can give it, you know, spark slash tell it, you know, the URI of your lake sale instance. Exactly. And then all you have to do is add the uh API equals connect. Yes. Uh to tell it to use the spark connect protocol. That's right. To talk to you guys. And because we're call already, we'll you guys already feed arrow. Well, spark connect itself as a protocol returns arrow data. Correct. That's the spark connect protocol works. Correct. So you guys implement the spark pro the spark connect protocol return arrow which means you get the zero copy no serialized d serialized super efficient column oriented the whole way end to end right right and so uh and you guys are also built on top of data fusion yes you apache data fusion which itself is arrow native also that's right that's right it was a part you told me again I learned a lot of new things part of the arrow project yeah originally data fusion was part of the arrow project and then it got the community grew large enough and the following got big enough that it was it spun off into its own top level Apache project right so and this is all open source which is which is awesome and so uh I just wanted to kind of uh segue a little bit so you mentioned that uh you guys started this Apache MacBY project and I think it's top of mind for many people how do you uh deal with AI right because a lot of development we do now is with AI and uh that is something you guys started I think a lot of folks I can't take credit there cuz I wasn't part of the one like I wasn't part of the ones who started it. Uh Apache Magpie grew out of uh the Airflow developer community, right? Um I I I got you because of people I know and being an ASF member uh I was asked to join the initial PMC for Magpie by the people who were organizing it. And so I that's how I became part of the project. But you know, but originally it spun out of the the tools and utilities that the Airflow maintainers were using were building for themselves. Yes. And and like AI skills, right, for to make it easier to maintain Airflow as a project in this age of LLMs and AI, right? You know, Mac Apache Magpie is is is a collection of skills and AI uh AI utilities for improving and making open-source maintainership better and easier, you know, triage, you know, triaging issues, triaging PRs, uh uh uh assisting with automated reviews, you know, and so on and so forth, making maintainers life easier, giving them time back. Yes. Yes. And I think I really like what you uh said yesterday that basically it gives you uh the confidence that somebody comm contributing to your project is is knowledgeable and thoughtful, right? Like it's fine to use AI. Like we're not anti- AI, but if you're using AI, you should be responsible for everything you send to the person. Exactly. Like like Magpie isn't like the skills of Magpie, you know, it requests confirmation before it ever posts anything on the pub on GitHub or anything. you know, you have to it says here is what I'm suggesting. Do you want me to post this? You know, it will it will also help with, you know, suggesting ways of helping to mentor new committers, new contributors, right? you know, while also helping you uh triage and al and align issues and PRs with labels and things like like so like the like like my workflow frequently lately with you know a arrows go implementation and iceberg go has been to use use magpie skills to do a you know a sweep of the PRs doing an initial review through and then I go through the reviews that magpie produced and then one by one go through those and review those myself before deciding okay yes I agree with that? Yes, we can pass that. No, change this and so on and so forth. So, you know, in order to actually have be able to keep up with the activity on on on the project, this kind of AI assistance for the review cycle is just invaluable for that way. Mhm. But the point there is that it also denotes anything anything posted, you know, hey, this was assisted by an AI but signed off by the maintainer. Exactly. Exactly. A a and the expectation is that you know any responses from the contributors you know are responsible for what they put right and that's that's the goal here is that it's it's it's as you said you're not anti- AAI but anything you're anything your LLM posts or sends a PR or what code writes you're responsible because your name's going to be on it right exactly exactly it's like they put on the bottles of fine scotch drink responsibly like usually responsibly Right. It's up to you. We're not prohibiting you from driving after one shot, but like it's like so like magpie is part of is part of the responsible AI committee at the ASF. Yes. Yes. You know that that's where many many of the people that are part of that that are the PMC and initial group for Apache Magpie are the responsible AI committee or at least related to people working on the responsible AI committee with the AASF you know which are which we are doing you know there's lots of suggestions and recommendations on you know policy around it and things like that just because at this point you can't just blanket ban AI no contributions it doesn't work it doesn't work you know And many open source projects have taken that tact and that's that's up to them to do. Yeah, it's they going to be condemned to be a niche because industrial development is I don't necessarily not necessarily on that one but but but for me the way I see it is that you you need the barrier to entry to be as low as possible because open source projects perpetually need more contributors and need more maintainers to be part of the projects. Mhm. And it's just a perpetual problem that open source software just does not have enough people Yes. on any given project. And so if the first way that someone engages with your project is to have an LLM help them file a PR. Mhm. I don't want to be a maintainer that just immediately goes, "Nope, that's AIGN, you know, not not look at it, not create AI, go away." Like I I don't want to do that. I want to encourage that person. I want to see are they someone who's going to engage with the project, right? You know, above and beyond being a meat proxy for their AI, right? Um and if they're willing to engage with the project and they're mentorable and you know, like AI, it's a new term I learned from you also. Meet proxy. It's a good one. Yeah. They they use they [laughter] used it in the keynote yesterday. Yeah. And I I heard about it and I've heard it a couple other places. Yes. Yes. You know, but but like like AI is LLMs and AI are a tool like anything else. And as developers are are we should be using as many tools as we can. Like I I've had a conversation with people where you know half of my career my job has been to automate my job, right? And that's going to be the same for most developers. Like you're the it's a neverending search for software engineers to automate as much as we possibly yourself out of the job. That's always and and that's the goal. That's the goal cuz cuz then you can move on to something else. That's right. And so like to to blanket be against AI is kind of weird in that scenario because it's just kind of the epitome of what we've been doing anyways in terms of trying to automate ourselves. It's just next step. It's next step. It's but as any but with but as any tool you have to know how to use it. You have to do it responsibly. You have to use it correctly. You know, if you're just vibe coding and that's all you're doing, cool. It's great for a prototype. Mhm. But if you're doing anything in production, you better be at least glancing at the code. You better but you better also be understanding what's going on and what it's doing. Mhm. So that you can, you know, steer your LLM at a minimum. Yes. You know, you the knowledge and experience of developers is still important. Yes. You have to show you have critical thinking capacity. So because because they're just going to do because the AI just going to do stupid things anyways. Exactly. Exactly. You need to be in the you show some experience and driving it. Yeah. So like going back to ADBC and so I really kind of I'm thinking right like like I mean I think we met at uh category VC and Wes was there and like sale and other like I see naturally self- selecting companies emerging who build fast AI infra in you know in rust and native technologies like ADBC in my mind it's like the new kids on the block like it's a new crop of tools which links this in like obviously arrow and data fusion and you know a bunch of Rust kind of systems. uh to me it's kind of all part of this and you are in the position where you can basically put ADBC into multiple systems right also your talk was about duck DB which is another cool kids database like people uh a a a uh so columner we we put out um a community extension for for duct DB for ADBC right that was uh that was it was it was done by uh a a PhD student that was an inter that was interning at columner there Sam March uh who built out the extension, right? And and we contributed to the community so that you can use ADBC drivers through duct DB to communicate with whatever system you want. And and the benefit there is because you know when you think about it, if you're using ductb locally, right, to do your analytics stuff, right? And then there's plenty of extensions people have for letting DuckDB communicate with external systems in efficient ways to talk to, you know, to fetch the data, right? And and that's the thing, analytics, the analytical part, this the SQL, that's the easy part. The hard part is getting the data over here, right? And so you need to have an efficient way to actually retrieve that data. Mhm. Whether you're going through ductb or you're going through a script yourself. Mhm. And so the ADBC extension of DDB lets you be able to leverage that column oriented zero copy ADBC arrow based interaction with any ADBC driver you want. You can attach to it. You can send the queries and pull the results set in and duct DB natively supports arrow anyways and so everything is is exceptionally efficient and super zero copy and fast. Yes. And duct DB makes it super easy to load extensions, right? So this is all becoming very natural. Exactly. I mean it's a community extension. You do, you know, install ADBC from community. That's right. Done. ADBC. Done. That's right. And and then you know and then you can easily mix local data with your remote systems. Mhm. Efficiently. Right. So and then you also mentioned that you're going to DBT summit and again from you learn that DBT is basing its driver strategy on ADBC DBT core v2 uh their data connectors and their and their their database connectors and their uh connect adapters are built on ADBC for their all their connections uh and DBT fusion is also using ADBC right so what I'm I'm just thinking right like uh we want AI infra it's it's about data AI agents are only as as the data that you feed to them and only as fast as you can feed it to them and as fast as you and everybody everybody wants to them to be you know fast and and correct and you know feed you know a lot of data to them and so it seems to me that uh you know if you're a good engineer building uh this AI in front of the future so you need to focus on the set of technologies so what what would would be your advice to like companies implementing distributed systems uh for AI so I think we mentioned some of pieces like you know data fusion like how can we get to the point where agents exchange data through a fast bus instead of like derializing things into JSON and then I mean because MCP is all built on JSON right like why would we spit out JSON feed it to something deserialize and serialize it again because people like human readable the human readable aspect of it you the other I mean the thing is like people will talk people will talk about how you know you can't you know you can't efficiently send huge amounts of data through MCP Right? But the reason why you can't just cuz it's JSON, right? You know, if you if it was a binary protocol, right, you could do it. It could be much more efficient, right? And but MCP is JSON. So like that we need a binary protocol for agent communication like that. That's right. That's right. You know, and and that's really going to be the next big thing I think is going to be someone pump, you know, standardizing a binary protocol for what MCP does, right? But the other big thing is also getting more and more of the vendors of your database vendors, your data source vendors to actually support returning arrow data, right? You know, you have so many systems send us columnar data. You have so many you have so many systems that are columnar data systems. Yes. That only return JSON, right? Trino Presto, you know, systems that only return row oriented data, right, from an analytical column oriented comput engine, right? and and you know uh uh uh uh uh uh you know and and so like it would be great to see more vendors and more systems and more applica more applications support arrow in and arrow out. Yeah, that's right. And even many of the systems that support arrow output still don't support arrow input yet. Mhm. Because because agents we're talking about aren't just about fetching the data so the agent can process it. Right. The agent also produces data. That's right. And you're going to want to write stuff back. That's right. An agent can be just a filter like it's a Unix pipe. It accepts error data and it sends data. It's crazy that the most efficient way we have right now in most systems, most like cloud systems to ingest data is to write it to parquet files on object storage and then run SQL queries to copy it into the data. Yes. because that system it'll you know bigquery and snowflake will output will output arrow but we even though ABC has a bulk ingestion the implementation for those drivers of the bulk ingest is to take the stream of arrow data and write it to parquet files on object storage and feed them through and then call copy into queries that's crazy because the system themselves don't have a native here's a stream of arrow data, right? Insert that into my table. Right. Right. Interesting. Even though I can do select star and get the entire table as arrow out. Zero out. Okay. Now, this I think this is great. Uh I think it it's probably not enough people asked for it or there was not enough push. But again, like this is I think the power of community if if we create enough kind of momentum uh behind the idea that we shall all interchange data through a fast bus behind all this AI system. [laughter] an open standard use an open standard then you know hopefully uh you know others dominoes will fall and uh we'll have a faster and and for those who like we mentioned arrow a bunch of times and I said you know inmemory format but for those who aren't familiar with apache arrow in general highly recommend you know look into you know arrow.apache.org The the important piece here is the fact that you're going to find that Arrow is and has been the underpinning of pretty much the entire data ecosystem for a while now. Arrow is 10 years old this year. Yes. And it's still not that well known even though you know it I will almost guarantee that any data stack you have arrows in there somewhere right because Polers is built on top of arrow right pandas uses arrow and has and has the pi arrow back end in v3 now. Yes, duct DB input takes in and outputs arrow. Yep. Snowflake outputs arrow. BigQuery outputs arrow. You know, PowerBI uses ADBC drivers now. Sale like sales is built on it's built on data fusion. Data fusion is entirely arrow. Correct. You know, like like like Arrow is the is the backbone of what makes all of these tools and utilities work together seamlessly and efficiently. Even though you're not dealing with Arrow directly yourselves. Yes. You know, if you're the one building the tools, then you use Arrow directly. If you're not the one building the tools, you aren't necessarily familiar with Arrow because you're you're just benefiting from the fact that all of these other things use Arrow to talk to each other. Right. Right. And so I think the agentic movement is kind of several steps removed. People talk about agents kind of communicating and then they do that they think chat, but the underlying data should be transferred through binary uh fast binary channel. It just made me think maybe you know MCP DBC is the next thing you [laughter] guys you can tackle on right like how hard can it be you know you have pie you have all the reviewing tools arrow MCP yeah you know that that you know you heard it first if I had if I had the time [laughter] so so hey someone can contribute it exactly contribute it well thank you very much Matt that was illuminating and you have a bunch uh things coming up for ADBC so huge you know kudos to you and thanks for sharing with the community and we See you around community. Thank you. Thank you. All right.