Release Notes · September 2026
Prefill/decode separation and multi-worker replica scaling, resumable cache management, multimodal embedding and rerank, plus Gemma-4, MonkeyOCR and a wave of new models — with Python 3.14 support.
What's New
Prefill/decode separation, native vLLM multiprocessing executor routing, and multi-worker replica scaling.
Resumable cache management, a download-only flow, launching straight from cached models, and oci:// model URIs.
Multimodal embedding and rerank, alongside Gemma-4, MonkeyOCR, dots ocr, MiniCPM5-2B and Fish Audio.
Python 3.14 support, a zh-TW locale, TTS streaming playback and reusable launch history.
Open Source
Installation
pip
pip install 'xinference==3.4.0'
Docker
docker pull xprobe/xinference:latest