Pricing

Most teams start on the Model API and move to dedicated compute as usage grows.
Every option runs the same platform and the same API.

Model API

Frontier open models by the token. Nothing to host.

For teams that want to ship on open models today.

Get API access

Cloud deployment

Dedicated GPUs in our cloud, fully managed.

For steady workloads that outgrow per-token billing.

Contact sales

Private deployment

The full platform inside your own perimeter.

For regulated teams: finance, health, government.

Contact sales

Xinference + Xagent bundle

Inference plus agents, one stack, one vendor.

For teams building agents on their own models.

Contact sales

Model API

Frontier open models served from Australia, priced below big-cloud rates for the same models. Billed per token.

For teams that want to ship on open models today.

Model What it's for Price
GLM 5.3 Agents, reasoning and coding Contact sales Enquire
DeepSeek V4 Long-context work and RAG Contact sales Enquire
Gemma 4 High volume at the lowest cost Contact sales Enquire
Xinference serves 300+ open models. Contact sales for more options.

Cloud deployment

Dedicated GPU instances in Xinference Cloud, fully managed, with zero infrastructure to run. Volume discounts available.

For steady workloads that outgrow per-token billing.

GPU Memory Price
H200 141 GiB VRAM Contact sales
H100 80 GiB VRAM Contact sales
B200 180 GiB VRAM Contact sales
A100 80 GiB VRAM Contact sales
RTX6000 Pro 96 GiB VRAM Contact sales

Xinference is NVIDIA optimised yet purpose-built to be hardware agnostic. Enquire with sales if other GPUs are required.

NVIDIA Inception program memberMember of the NVIDIA Inception program

Xinference on AWS Marketplace

Deploy through the AWS account you already have. Billing counts toward your AWS spend commitments.

Contact sales

Private deployment

Deploy on your own cloud, your data centre, or fully offline environments, with full controls and dedicated support. No data crosses your perimeter.

For regulated teams: finance, health, government.

Xinference + Xagent bundle

Private deployment plus the Xagent agent platform. From $10K a month for 500 users, hosted in Australia.

For teams building agents on their own models.

Community vs Enterprise

Community and Enterprise run the same Apache-2.0 engine.
Enterprise hardens it for production: up to 2x throughput, broader hardware, failover and named support.

Community Enterprise
Engine Apache-2.0 inference engine The same engine, plus commercial extensions
Hardware support NVIDIA only (validated) NVIDIA, AMD, Intel, CPU
Scaling Single cluster Multi-cluster, multi-region, HA
High availability Manual restart on failure Automatic failover, production SLA
Performance tuning Default engine settings Tuned per workload by Xinference engineers, up to 2x throughput
Security API keys RBAC, immutable audit logs, SSO via Google Workspace (SAML on the roadmap)
Compliance Roll your own Compliance support for regulated deployments
Data residency Manual Regional deployment, DPA, zero data leaves your environment
Private offline Community-supported Certified offline Helm chart, dedicated support
Support GitHub community Named contact, priority response, quarterly engineering reviews

Pricing questions

How is the Model API billed?

Per token, per model. Contact sales for current rates and volume pricing.

Can we move between options later?

Yes. Every option runs the same OpenAI-compatible API, so a move is a base URL change, not a rewrite.

Which models can we run?

300+ open models including GLM, DeepSeek, Gemma, Qwen and Llama, plus your own fine-tunes. Contact sales for the current catalogue.

Is our data used for training?

No. No training on your data, no retention by default, processed in Australian data centres or on your own hardware.

What does Enterprise add over Community?

Up to 2x throughput, broader hardware, automatic failover and named support. See the comparison table above.

Can we pay in AUD?

Yes. AUD invoicing is standard for Australian customers.

How do we start?

Ask for a cost assessment: we model your current workload side by side, and you see the savings before committing.

Start on the API. Scale on your terms.

Contact sales