talk · community record
All About Evaluating LLM Applications // Shahul Es // #179
Deep dive into evaluation of open source models: debugging, troubleshooting, the limits of public benchmarks, the importance of custom data distributions, and the role of fine-tuning. Also on YouTube at youtube.com/watch?v=LOpv3vQeLxU.
01
Connections
1 relationship