Scale By The Bay 2020: Jean-Ives Stephan, Cloud-Native Apache Spark: why & how to migrate your Spark
ai.bythebay.io Nov 2025, Oakland, full-stack AI conference Title: Cloud-Native Apache Spark: why and how to migrate your Spark pipelines to Kubernetes Apache Spark can run on top of Kubernetes (as opposed to Hadoop YARN or Standalone mode) since Spark versions 2.3 (2018). In the past two years, the support for running Spark on Kubernetes has grown a lot, and a lot of companies have adopted it -- in fact, Spark-on-Kubernetes will be officially considered "production ready" with the upcoming release of Spark 3.1. In this talk, we will go over the main reasons why many companies decide to adopt Spark-on-Kubernetes, and our best practices for making Spark on Kubernetes reliable and performant at scale. No prior knowledge of Spark or Kubernetes is required, but you should expect a technical session heavy with code-examples and real-life tips to help you productionize Spark on Kubernetes. Jean-Ives Stephan Data Mechanics Co-Founder & CEO JY is the co-founder of Data Mechanics, a cloud-native Spark platform making Spark easy-to-use and cost-effective for data engineers. Their platform is deployed on a Kubernetes cluster inside their customers cloud account (AWS, GCP, and Azure are supported). Prior to Data Mechanics, JY was a software engineer at Databricks. JY is passionate about serverless architectures and making data infrastructure 10x more easy-to-use and efficient through the use of automation.