video · community record
DevReal: Optimizing LLMs for Cost-Efficient Deployment with vLLM - Michael Goin
Cutting-edge compression techniques and advanced inference system optimizations that enable fast LLM performance on hardware of choice. Practical strategies and tools enterprises use to scale deployments while minimizing costs.