talk · community record
Optimizing LLMs for Cost-Efficient Deployment with vLLM
Cutting-edge compression techniques and advanced inference system optimizations that enable fast LLM performance on hardware of choice. Practical strategies and tools enterprises use to scale deployments while minimizing costs.
01
Connections
6 relationships
aboutvLLMproject ↗affiliated withNeural Magiccompany ↗documented byBay Area AI 20250123photo ↗presented · incomingMichael Goinperson ↗presented at[registration required on lu.ma/_ai] Full Stack OSS AI: Chips to Apps!event ↗recorded asDevReal: Optimizing LLMs for Cost-Efficient Deployment with vLLM - Michael Goinvideo ↗