One control plane,
from app to GPU.
One integration that keeps working as models, hardware and traffic change. Deploy, route and govern open models across cloud and on-prem infrastructure, with one API for your applications.
How it fits
Your applications
LangChain
LlamaIndex
Dify
Your code
one OpenAI-compatible call

by cost, latency, availability
Models and GPUs
Qwen
DeepSeek
Kimi
GLM
+300 more
What you get
Model repository
300+ open models, embedding and multimodal, plus your own fine-tunes, deployed from one catalogue.
Routing
Our proprietary smart routing engine optimises cost and performance. We can also pair you with an engineer to tune it for your workload.
Deployment options
Our cloud, your cloud, or your data centre, including fully offline environments.
Governance
Role-based access at organisation and resource level, audit logs, SSO via Google Workspace, and live monitoring of TTFT and TPOT.
Engines under the hood: vLLM · SGLang · MLX · Transformers · one OpenAI-compatible REST API with streaming.
Running 50+ AI use cases in production for Yum! Brands.
Where it runs
Xinference Cloud
Your own dedicated environment in Australian data centres.
Your infrastructure
Install into your data centre or your cloud account, behind your own perimeter. Built for regulated workloads.
AWS Marketplace
Deploy through AWS Marketplace and keep billing on the AWS account you already have.
Every model type
LLMs: chat, code and reasoning with long-context streaming.
Embeddings and search: dense retrieval for RAG and semantic search.
Multi-model pipelines: embedding, rerank and LLM on one runtime.
Transcription: real-time streaming speech to text.
Text to speech: low-latency voices for calls and agents.
Image generation: FLUX, SDXL and SD3.