Bay.Area.AI: Efficiently serving LLMs at scale, Nick Hill
ai.bythebay.io Nov 2025, Oakland, full-stack AI conference Efficiently serving LLMs at scale, Nick Hill, IBM In this talk I will discuss challenges of serving LLMs efficiently in highly concurrent, multi-user contexts, and some of the optimizations unique to these kinds of models that have emerged over the the last year. These include "continuous batching" of heterogeneous requests. I'll dig into the implementations which involve careful manipulation of tensors with PyTorch. Nick Hill, IBM. Nick is a Senior Research Engineer focused on scalable serving of large language models. He previously led the architecture and development of distributed machine learning infrastructure supporting key IBM AI cloud products and services including Watson Assistant, Watson Discovery and Watson Natural Language Understanding. He designed and implemented the Model-Mesh serving framework that supports hundreds of thousands of models, now a key component of the KServe open source project. He is also an author of and contributor to other open source projects.