A 10-week, 30-minutes-a-day roadmap for LLM inference serving and optimization. vLLM, SGLang, quantization, speculative decoding, benchmarking.
331stars29forksHTML
learning-resourcesllmllm-inferencemlopsroadmapsglangvllm
Real data pulled from GitHub this week. The author's original repo lives upstream.
View on GitHub