๐
New Models
NEWMiniCPM5-1B
Compact language model
MiniCPM-V-4.6
Vision-language model series
jina-embeddings-v5
Embedding model series
Tencent Hy-MT2
Machine translation series
VoxCPM2
Speech / audio model
PaddleOCR-VL-1.6
OCR vision-language model
โจ
New Features
๐จ
Revamped Web UI
Continuous improvements to the new web interface
๐
Password Management
Local users can now change their password
๐งช
Embedding Test Page
New Gradio-based embedding testing interface
โ๏ธ
GLM-5 Tool Parser
vLLM integration for GLM-5 tool calling
๐ฒ
Serving Benchmark
New random serving benchmark workload
๐ก๏ธ
Model Loading State Machine
Prevents traffic from reaching models that are still loading
๐
Enhancements
- Updated model JSON configs (LLM / Embedding / Audio)
- CI migration to GitHub Hosted Runners
- Log center interface optimization
- Enhanced monitoring rules with new alert policies
- Improved model recovery and replica management
๐
Bug Fixes
Supervisor restart worker registration recovery
Worker Monitor startup timing fix
OOM pool recovery fix
Flexible model registration fix
Docker CPU image kernel version issue
FastAPI version compatibility fix
vLLM single-GPU EngineCore scheduling fix
GPU orphan process cleanup fix
Rerank llama.cpp parameter passing fix
Anthropic system prompt handling fix
Qwen3-Next Hybrid KV Cache auto-patch fix
Qwen3-ASR temperature parameter compatibility
๐
Documentation
- Updated docs logo and favicon
- Updated Read the Docs build image
- Removed Zhihu links from English docs
๐ข
Enterprise Edition
ENTERPRISE
๐
AutoTune
New auto-tuning capability for optimal performance
๐ก
Stability Fixes
Numerous stability fixes and performance optimizations
Get Started
๐ฆ pip
pip install 'xinference==2.11.0'
๐ณ Docker
Pull the latest image or run pip install inside your container.