Webinar, 15 Oct: Who controls your enterprise AI? Join us live →

Book · Free download

Inference at scale

Engineering lessons from a decade of distributed systems. Why statically configured systems fail at scale, and which runtime decisions set the cost of serving models.

By Chris Qin, Co-founder and CEO, Xinference

Cover of Inference at Scale

What’s inside

The research behind it

What the Xorbits paper (ICDE 2024) showed about planning at runtime instead of in advance.

From dataframes to tokens

How the same ideas apply to LLM inference: batching, KV cache and engine choice.

A production checklist

Questions for the technical leader and checks for the engineer running the system.

A glossary

Plain definitions of the serving terms used in the book.

PDF, 18 pages.

Get the PDF