talk · community record
Episode 40: What Every LLM Developer Needs to Know About GPUs
Charles Frye gives a practitioner's tour of GPU architecture for LLM developers - memory hierarchy, throughput versus latency, and how hardware constraints shape inference cost. Recorded as a livestream, also at https://www.youtube.com/watch?v=INryb8Hjk3c
01
Connections
1 relationship