Book · Free download
Inference at scale
Engineering lessons from a decade of distributed systems. Why statically configured systems fail at scale, and which runtime decisions set the cost of serving models.
By Chris Qin, Co-founder and CEO, Xinference
What’s inside
The research behind it
What the Xorbits paper (ICDE 2024) showed about planning at runtime instead of in advance.
From dataframes to tokens
How the same ideas apply to LLM inference: batching, KV cache and engine choice.
A production checklist
Questions for the technical leader and checks for the engineer running the system.
A glossary
Plain definitions of the serving terms used in the book.