Alexy Khrabrov
Community builder · FunctionalTV
Alexy Khrabrov's graph
Centered on this profile · drag to explore
Community builder · FunctionalTV
Centered on this profile · drag to explore
Showing 1 of 1
A database system that can synchronize all or a part of its contents over a limited bandwidth link is described. The lowest layer of the system, the bedrock layer, implements a transactional block store. On top of this is a B+-tree that can efficiently compute a digest (hash) of the records within any range of key values in O(log n) time. The top level is a communication protocol that directs the synchronization process such that minimization of bits communicated, rounds of communication, and local computation are simultaneously addressed.
Showing 12 of 12
With the growing availability of data within various scientific domains, generative models hold enormous potential to accelerate scientific discovery. They harness powerful representations learned from datasets to speed up the formulation of novel hypotheses with the potential to impact material discovery broadly. We present the Generative Toolkit for Scientific Discovery (GT4SD). This extensible open-source library enables scientists, developers, and researchers to train and use state-of-the-art generative models to accelerate scientific discovery focused on organic material design.
Showing 24 of 547
Vlad Luzin, co-founder and CTO of BAND, with Alexy Khrabrov on letting agents talk to each other. Connecting agents is a distributed system of microservices and most of the work is plumbing, so BAND supplies the primitives: a registry in which agents discover each other and work as a group but cannot reach outside it without consent, an identity for each agent, and token and credential propagation so one person's agent acts in that person's scope. The registry is a deterministic data model rather than a compile-time decision, and it belongs to a human -- every agent, coding assistants included, has an identity and an owner, and the owner decides who joins.
Kranti Parisa, co-founder of LaserData, with Alexy Khrabrov on what they call the AI Infra 2.0 movement. The argument: agentic infrastructure is still being built on legacy stacks, and every hop across a JVM boundary forces serialization, so Amdahl's law caps whatever parallelism can recover -- efforts to accelerate JVM systems reach single-digit speedups where Rust systems reach ten to a hundred times. LaserData's streaming engine, built on Apache Iggy, is Rust, with millisecond P99 response times and three to four times the throughput. Parisa argues that context propagation at millisecond latency matters more in production than inference cost, because an agent working from stale context wastes compute, multiplies subsystem calls and pollutes observability data.
Forty-two seconds from the floor of the WeAreDevelopers World Congress North America, the series' North America debut, at the San Jose McEnery Convention Center with more than ten thousand attendees. Alexy Khrabrov's read on the hall: it is all AI, all agents, all AI thought leadership, and how you cope with that much change -- plus the startups in uniforms selling their tooling, which he finds instructive, and a great many good people in the audience.
Shadaj Laddad and Chris Wensel, two of the SF Systems Club's organisers, interviewed by Alexy Khrabrov at the meetup. Alexy's history with Shadaj runs back to 2011, when Shadaj gave his first talk as a child at the Scala meetup Alexy had just started in the Bay Area; both his parents were long-time speakers there too, which Alexy and Shadaj agree makes it a family business. Shadaj works at the interface of functional programming and distributed systems, and now AI, and leads the Hydro team at AWS. Chris Wensel is the creator of Cascading. Links Chris Wensel shared afterwards, to share freely: Retrofit (https://retrofit.sh), intrastate (https://github.com/cwensel/intrastate) and arcaneum (https://github.com/cwensel/arcaneum).
Alexy Khrabrov's note from the floor of the SF Systems Club meetup of 24 September 2026, hosted by LatchBio. The club is organised by Shadaj Laddad -- a long-time speaker at the Scala, Scale By the Bay and AI By the Bay meetups Alexy ran, who took his PhD at UC Berkeley and is now at AWS -- and Alexy's verdict on the evening is that it was a fantastic meetup.
Arun Sharma, founder of LadybugDB and Ladybug Memory, interviewed by Alexy Khrabrov at the Rows & Columns Summit. LadybugDB continues Kuzu, the embedded graph database developed at the University of Waterloo whose team was acquired by Apple: a well-regarded codebase with published research behind it and no community, which Sharma has spent eleven months building one around, with Ladybug Memory supporting the work and aimed at agentic memory. Before that, five years at Google and then Facebook, where he built a graph indexing system sitting beside the world's largest MySQL cluster, listening to its write-ahead log and built on RocksDB as its very first user, and in 2018 prototyped what today looks like Amazon DSQL.
Alexy Khrabrov's recording from the floor of the Rows & Columns Summit, at the Contemporary Jewish Museum in San Francisco on 22 September 2026 -- a practitioner-first, single-track conference on the architecture question that will not go away: OLTP and OLAP, together or apart. Andy Pavlo opened the day, and Hannes Muehleisen, co-founder and creator of DuckDB, argued that nobody knows what OLTP is and that DuckDB is moving to the middle. Alexy's own frame is fast access to data in real time set against transactional access, and why merging the two is becoming paramount for agents of every kind. He closes by saying the interviews he recorded at the summit will appear on struct.fm alongside the back catalogue.
Alexy Khrabrov's notes from PyData Amsterdam 2026, recorded after the trip. The LakeSail team met in person for the first time around Shehab Amin and Santosh Pingale's talk on Sail at Adyen, where Spark jobs that ran out of memory or never finished now run on a Rust engine behind the same PySpark API. Why the warehouse at the NDSM Loods forced Alexy to record his own audio. Christophe Blefari's keynote on the history of analytics, and why nao, his open-source analytics agent, is the user every lakehouse builder should design for. Ritchie Vink tracing the engineering arc of Polars from a better pandas to a company, and why Rust-native distributed data systems strengthen the whole ecosystem. Matt Topol on ADBC adoption, ADBC for DuckDB, and the ADBC Spark driver that lets Sail plug in anywhere Spark Connect is spoken; and Apache Magpie's tools for maintainers facing AI-generated pull requests. Plus the QueryGraph stack Alexy has been building on Sail: Grust graphs, TypeSec policies, Lake
Alexy Khrabrov's recollection of the first Rust AI meetup in Europe, held at Adyen in Amsterdam on September 9, 2026, the night before PyData. How Adyen, processing a trillion dollars a year, found Sail on its own when Spark jobs ran out of memory, and why a meetup organizer with thousands of people in his Bay Area ecosystem has to start over in a new city: co-hosting with AI Foundry, restoring the 2018 Rethink Trust attendee list with Claude and Codex from Gmail and Drive, and holding the event through a nationwide Dutch transportation strike. Zemin Piao's talk on evolving Spark at Adyen with a single URL change; Shehab Amin and Heran Lin, LakeSail's co-founders, presenting together for the first time; Robin Everaars on digital sovereignty and putting Sail to work for the Dutch Ministry of Defence. Why the meetup is about fast, AI-native foundations under Python APIs, whether Rust, C++, Zig, or OCaml, and never the JVM: Amdahl's law always wins.
Alexy Khrabrov talks with Matt Topol, co-founder of Columnar and PMC member of Apache Arrow, Apache Iceberg, and the new Apache Magpie, about ADBC, the Arrow-native replacement for ODBC and JDBC that keeps data columnar end to end. They discuss dbc, Columnar's package manager for signed ADBC driver binaries, the ADBC community extension for DuckDB, Spark Connect and Sail returning Arrow natively, dbt building its adapters on ADBC, why agent protocols need a binary channel rather than JSON, and how Apache Magpie's skills help open-source maintainers use AI responsibly.
Alexy Khrabrov talks with Christophe Blefari, co-founder of nao Labs, after his PyData Amsterdam 2026 keynote on the history of analytics from the warehouse to the lakehouse and today's agentic systems, including a live demo of talking to data in DuckDB. They discuss what a semantic layer should be, with unambiguous, human-readable definitions of metrics and dimensions rather than a pile of SQL queries; a two-layer approach where an agent falls back from the strict semantic layer to broader context; and nao, an open-source analytics agent that lets everyone in a company chat with its data while data people act as context engineers, with bring-your-own model and database.
Alexy Khrabrov talks with Ritchie Vink, founder of Polars, at PyData Amsterdam 2026, Ritchie's home game. They discuss why Polars was written in Rust six years ago and why Rust's compile-time guarantees now make it a strong language for AI-assisted coding; how database research, with lazy evaluation, a query optimizer, a consistent relational data model, and strict column types, shaped Polars in contrast to pandas; Polars as a Python-first library that catches type errors before a query runs; a growing focus on SQL for agents; and the goal of being the fastest engine at any scale, including distributed.
Alexy Khrabrov's PyData Amsterdam 2026 lightning talk on the QueryGraph stack, an open-source layer built on Sail, LakeSail's Rust implementation of Spark: pip install pysail, Python UDFs running from Rust without crossing the JVM boundary, and a partner ecosystem of Rust and Python startups.
The lightning talks of PyData Amsterdam 2026, recorded on 10 September at NDSM Loods. Carlos Morales of Portima on the token scarcity problem: rate limits per minute when every user opens their email analyzer at nine in the morning. Daniel Pacheco on the stable marriage problem and its algorithm. Jeroen on why only one percent is scared of AGI. Muhammad Chenariyan Nakhaee on how Python made his addiction to vintage cameras worse. Marijn Markus on data everywhere: a career from statistics through big data and data science to AI. Riaan Zoetmulder on how to tame your social media, starting from Bo Burnham's Welcome to the Internet. Alexy Khrabrov closes with the QueryGraph stack on Sail, published separately as its own recording.
Alexy Khrabrov talks with Adam de Delva of DTR (Developer Technology Research) in New York after the Apache Spark NYC meetup at Datadog. They discuss growing a network of open-source communities from 30,000 to over three million members, with new communities in Ghana and Latin America; open-source sustainability and getting companies to reinvest in the commons; digital public infrastructure and sovereign AI, from mapping workloads onto neoclouds to scheduling containers securely in trustless environments with the Linux Foundation Decentralized Trust and Hyperledger communities; and the UOR (Universal Object Reference) Foundation, where specs, protocols, and standards for sovereign AI infrastructure are built.
Alexy Khrabrov talks with Yarden Wolf, Data Engineer and AI Tech Lead at Wix, at the Apache Spark NYC meetup at Datadog. They discuss how agentic workflows change the work of Wix's software engineers, data scientists, and analysts while keeping a human in the loop for production alerts; Base, Wix's app-building platform, and Wix Harmony, which combines vibe coding with drag-and-drop editing; authentication and AI gateways for internal agents; specialized internal agents such as Airbot for Airflow alerts and a root-cause-analysis agent working together; and Wix Headless, which lets a coding agent build a site's backend on Wix.
Alexy Khrabrov talks with Romain Priour of LangChain’s AI infrastructure team about deploying LangSmith for enterprise customers. They discuss Kubernetes, a hybrid SaaS and self-hosted model on AWS, customer-facing infrastructure, and expanding the platform across clouds.
Alexy Khrabrov talks with Hotdata co-founder and principal engineer Zac Farrell about building cloud databases for AI agents. They discuss Hotdata’s Rust and Apache DataFusion foundation, isolated database workspaces, SQL, performance engineering, and choosing Rust-native infrastructure over legacy stacks.
Alexy Khrabrov walks through the new LangChain office before the San Francisco Apache DataFusion Meetup and checks in with organizers and speakers. Divya Ranganathan, Emil Sadek, Shehab Amin, and Alexander Bianchi preview talks spanning ADBC, LakeSail, distributed query execution, and the growing DataFusion community.
Alexy Khrabrov talks with LangChain co-founder and CTO Ankush Gola about SmithDB, the Rust and Apache DataFusion-powered data layer behind LangSmith. They discuss agent-observability workloads, database extensibility, production performance, and where Rust fits into modern AI infrastructure.
Alexy Khrabrov presents the QueryGraph lakehouse stack: Sail, an agentic semantic layer, LakeCat governance, Typesec and TypeDID security, and the Grust graph API.
Shehab Amin discusses rebuilding the Spark-compatible lakehouse ecosystem in Rust with Sail, Apache Arrow, Apache DataFusion, Delta Lake, and Apache Iceberg REST.
Alexy Khrabrov opens Rust AI Begins, introduces the San Francisco Rust-and-AI community, and previews the evening’s projects and speakers.
Showing 7 of 7
Meetup photo album
Open gallery ↗Meetup photo album
Open gallery ↗Conference photo album
Open gallery ↗Conference photo album
Open gallery ↗Conference photo album
Open gallery ↗Conference photo album
Open gallery ↗Conference photo album
Open gallery ↗