Talk to an engineer
Product

Your models. Your infrastructure. Nothing leaves.

Deploy 300+ open models behind one OpenAI-compatible API, in your private cloud or on-prem. Nothing routes through a third party.

Standard bundle · Sovereign AI
LLM, Inference, AI Agents and Hosting
$10K/month
  • Unlimited AI agents
  • 500 users included
  • Dedicated private LLM, Australian hosted
Trusted by teams running AI in-house
Siemens
AIA
Yum! Brands
Everbright Securities
TFC OpticalComms
Berry Genomics
XW Bank
NVIDIA Inception Program
Runs the open models and tools you already use
gpt-ossDeepSeekQwenKimiLangChainLlamaIndexDify

“Xinference pools heterogeneous GPU resources in our private cloud, significantly reducing infrastructure cost for 6,000+ users while turning the latest open-source models into core productivity.”

ES
Everbright Securities
Financial services
Read the story →

Built around one idea.

Your models, on your infrastructure, at a lower and more predictable cost.

Nothing leaves your boundary.

Every model, prompt and embedding runs inside your own cloud account or your own data center. No third-party API ever sees your data.

Predictable, not metered.

Fixed infrastructure cost instead of per-token billing, with every dollar visible on one screen instead of a surprise usage invoice.

Faster than a closed API.

Up to 4x more models per GPU and 2 to 4x lower latency, so the same hardware serves far more production traffic.

Everything the platform does for you.

From single-command deploys to enterprise-grade clusters.

OpenAI-compatible API

Point your existing SDK at a new base URL. No rewrites, no new client libraries, no migration project.

300+ open models

gpt-oss, DeepSeek, Qwen, Kimi and more, ready to launch with one command from a single catalog.

NVIDIA GPU scheduling

NVIDIA GPUs, with support for non-NVIDIA accelerators, so capacity never sits idle.

Enterprise permissions and audit

Role-based access, audit logs and version history built into the control plane, not bolted on.

Private, hybrid or managed

Run in your own VPC, on-prem on your own hardware, or let Xinference manage the cluster for you.

Hot model swapping

Swap models without downtime, with multi-tenant isolation keeping every team and workload separate.

Xinference Enterprise

Better performance and enterprise-grade reliability.

Features
Standard
Enterprise
Hardware support
NVIDIA GPUs
NVIDIA GPUs, plus non-NVIDIA accelerators
Models per GPU
One model per GPU
Multiple models, higher utilisation
Management console
CLI and dashboard
RBAC, audit logs, unified console
Throughput
Baseline
Up to 4x more models per GPU

Read the Xinference documentation ↗

Model library

300+ open models, one catalog.

Pull from HuggingFace or push your own. New releases land day one.

gpt-oss 120B

An open-weight model built for self-hosting, one command to launch.

DeepSeek V4

Frontier reasoning and code generation, fully open weights.

Llama 3.3 Instruct

A general-purpose workhorse for chat, RAG and fine-tuning.

Qwen 3.5

Long-context multilingual chat with strong agentic tool use.

Kimi K2-Instruct

Large-scale reasoning with efficient mixture-of-experts routing.

GLM 4.7

Compact, efficient reasoning built for high-concurrency deployments.

Browse the model library

Proof, not promises.

Teams cutting cost and shipping models faster with Xinference.

~50%
faster first response, 500K+ daily requests on a regulated AI platform
AIA
65%
GPU utilisation across 50+ AI use cases and 1.52M requests/day
Yum!
JL

We consolidated model serving onto Xinference and stopped babysitting the stack ourselves.

Jason L.
VP of Engineering
TFC OpticalComms
Deploy

Your infrastructure, your control.

Self-hosted by default, in your environment end to end.

Fully managed

Everything to run sovereign AI: $10K per month.

One standard bundle with a dedicated private LLM, hosted in Australia, ready to go.

  • Unlimited AI agents
  • 500 users included
  • Dedicated private LLM, Australian hosted
Your environment

On your own infrastructure, cloud or on-premises.

Tailored deployments for teams that need Xinference inside their own perimeter.

  • No data crosses your perimeter
  • Tailored to your requirements
  • Best for regulated data: finance, health, government
Talk to an engineer
Built for enterprise IT and security

Governance your compliance team will sign off on.

Audit logs
Every request, deployment and configuration change logged and exportable.
Role-based access
Fine-grained permissions by team, environment and model.
Model version governance
Pin, roll back and approve model versions before they reach production.
Data residency
Choose AU or SG for managed deployments, or keep everything on your own hardware.
SSO and SAML
Provision and de-provision access through your existing identity provider.

Closed API vs. self-hosted open inference.

Same models, same requests, a very different bill.

Closed-source API

Billed per token, cost scales directly with usage. No visibility into spend until the invoice arrives.

Self-hosted with Xinference

Fixed infrastructure cost. The same GPUs serve every model and every team, with full utilisation visibility.

UPTO70%*
lower cost than closed-source APIs

*Depending on your workload, model choice and deployment path.

Frequently Asked Questions

Everything you need to know about Xinference and how it fits into your AI stack.

Nothing leaves.
Everything scales.

Deploy your first model in one command, on infrastructure you control.