Webinar, 15 Oct: Who controls your enterprise AI? Join us live →

Reading list · Free download

The Adaptive Inference Reader

An annotated reading list of 14 research papers on efficient LLM inference, with a short note on why each one matters in production.

Cover of The Adaptive Inference Reader

What’s inside

Scheduling foundations

Where runtime planning starts, from Xorbits to FlashAttention.

Memory that adapts to the request

PagedAttention and SGLang, and what they change about serving.

Spending compute where it counts

Speculative decoding, Medusa and early exit.

Precision and serving phases

Quantization with GPTQ and AWQ, and splitting prefill from decode.

PDF, 10 pages.

Get the PDF