talk · community record
SBTB 2023: Chris Matteson, Cost and Performance Optimization of LLM Inferencing.
ai.bythebay.io Nov 2025, Oakland, full-stack AI conference This session provides an overview of the architectures and associated considerations when running AI models and inferencing (i.e., together, “AI Applications”) for various use cases, delving more deeply into the use case of most interest to enterprises: event-driven, user-interactive AI applications. Finally, we will dive into common pricing models that arise from these architectures in light of their respective abilities to control costs based on the architectures. More details available here: https://www.scale.bythebay.io/post/chris-matteson-cost-and-performance-optimization-of-llm-inferencing