Nothing leaves your boundary.
Every model, prompt and embedding runs inside your own cloud account or your own data center. No third-party API ever sees your data.
Deploy 300+ open models behind one OpenAI-compatible API, in your private cloud or on-prem. Nothing routes through a third party.



“Xinference pools heterogeneous GPU resources in our private cloud, significantly reducing infrastructure cost for 6,000+ users while turning the latest open-source models into core productivity.”
Your models, on your infrastructure, at a lower and more predictable cost.
Every model, prompt and embedding runs inside your own cloud account or your own data center. No third-party API ever sees your data.
Fixed infrastructure cost instead of per-token billing, with every dollar visible on one screen instead of a surprise usage invoice.
Up to 4x more models per GPU and 2 to 4x lower latency, so the same hardware serves far more production traffic.
From single-command deploys to enterprise-grade clusters.
Point your existing SDK at a new base URL. No rewrites, no new client libraries, no migration project.
gpt-oss, DeepSeek, Qwen, Kimi and more, ready to launch with one command from a single catalog.
NVIDIA GPUs, with support for non-NVIDIA accelerators, so capacity never sits idle.
Role-based access, audit logs and version history built into the control plane, not bolted on.
Run in your own VPC, on-prem on your own hardware, or let Xinference manage the cluster for you.
Swap models without downtime, with multi-tenant isolation keeping every team and workload separate.
Better performance and enterprise-grade reliability.
Read the Xinference documentation ↗
Pull from HuggingFace or push your own. New releases land day one.
An open-weight model built for self-hosting, one command to launch.
Frontier reasoning and code generation, fully open weights.
A general-purpose workhorse for chat, RAG and fine-tuning.
Long-context multilingual chat with strong agentic tool use.
Large-scale reasoning with efficient mixture-of-experts routing.

Compact, efficient reasoning built for high-concurrency deployments.
Teams cutting cost and shipping models faster with Xinference.
Self-hosted by default, in your environment end to end.
One standard bundle with a dedicated private LLM, hosted in Australia, ready to go.
Tailored deployments for teams that need Xinference inside their own perimeter.
Same models, same requests, a very different bill.
Billed per token, cost scales directly with usage. No visibility into spend until the invoice arrives.
Fixed infrastructure cost. The same GPUs serve every model and every team, with full utilisation visibility.
*Depending on your workload, model choice and deployment path.
Everything you need to know about Xinference and how it fits into your AI stack.
Deploy your first model in one command, on infrastructure you control.