talk · community record
Efficiently Serving LLMs at Scale
Challenges of serving LLMs efficiently in highly concurrent, multi-user contexts, including continuous batching of heterogeneous requests and careful manipulation of tensors with PyTorch.
Challenges of serving LLMs efficiently in highly concurrent, multi-user contexts, including continuous batching of heterogeneous requests and careful manipulation of tensors with PyTorch.