Reading list · Free download
The Adaptive Inference Reader
An annotated reading list of 14 research papers on efficient LLM inference, with a short note on why each one matters in production.
What’s inside
Scheduling foundations
Where runtime planning starts, from Xorbits to FlashAttention.
Memory that adapts to the request
PagedAttention and SGLang, and what they change about serving.
Spending compute where it counts
Speculative decoding, Medusa and early exit.
Precision and serving phases
Quantization with GPTQ and AWQ, and splitting prefill from decode.