SBTB23:Oleg Avdeëv& Riley Hun, Lesson learned from orchestrating large-scale GenAI, ML & data on k8s
ai.bythebay.io Nov 2025, Oakland, full-stack AI conference Metaflow, a Python framework for ML/AI infrastructure which was originally open-sourced by Netflix in 2019, has come a long way. From its AWS-native roots, it has expanded to support all major clouds and on-prem deployments powered by Kubernetes. Recently, Metaflow gained support for MPI-style parallel compute, distributed training with PyTorch and Ray, opening up many new use cases around GenAI which require clusters of even hundreds of GPUs. In this infrastructure-focused talk we share a number of lessons learned from a diverse set of data intensive workloads. This session should be informative for platform engineers, data scientists and ML engineers curious about ML infrastructure, and anyone interested in building real-world, production-grade ML/AI systems in general. More details available here: https://www.scale.bythebay.io/post/oleg-avde%C3%ABv-lesson-learned-from-orchestrating-large-scale-genai-ml-and-data-on-kubernetes