Xinference
Models Deployment
Australian Sovereign AI

Private models. Real savings. Your control.

LLM inference, AI agents and hosting on open models — run in Australia, at a predictable cost per token. Up to 70% cheaper than closed-source APIs.*

console.xinference.co/deployments
syd-1
Deploy model ×
glm-5.2
1M context
✓
kimi-k3
1M context
llama-4-scout
10M context
region: syd-1 Deploy
Deployments
4 models · private cluster · syd-1 5 models · private cluster · syd-1
Deploy model
Requests 24h
6.42M
GPU utilisation
78%
Data egress
0 B
Model
Status
Requests
p95
DeepSeek-V4-Flash
Running
1.24M
310 ms
Kimi-K3
Running
842K
480 ms
Qwen-3.8
Running
216K
390 ms
Gemma-4
Running
4.10M
120 ms
GLM-5.2 new
Provisioning Running
—
—
Get your private AI model

Switching takes two lines.

*Compared with closed-source API list pricing, based on customer-reported infrastructure comparisons for self-hosted open models. Actual savings vary by workload and utilisation.

deepseek-v4-flash kimi-k3 qwen-3.8 gemma-4 gpt-oss-120b llama-4-scout whisper-large-v3 bge-m3 glm-5.2 mistral-large-3 qwen3.8-coder
Trusted in production
NVIDIAINCEPTION PROGRAM
Siemens Yum! Brands AIA TFC OpticalComms XW Bank

“Xinference gives us one platform to run every model on infrastructure we fully control.”

Marcus Z. · Senior AI Infrastructure Lead, Siemens

“We consolidated model serving onto Xinference and stopped babysitting the stack ourselves.”

Jason L. · VP of Engineering, TFC OpticalComms
Drop-in replacement

Switching takes two lines.

If your product talks to OpenAI today, it can talk to Xinference tomorrow. Change one address and one key — everything else stays exactly as it is.

+2 lines · that’s it
 from openai import OpenAI
  
 client = OpenAI(
+    base_url="https://api.xinference.co/v1",
+    api_key="XINFERENCE_API_KEY",
 )
  
 # done — your app now runs on Australian soil.
Private by design

Your AI data stays yours.

Built for organisations that cannot send prompts to an offshore API and hope for the best.

No data offshore

Prompts and outputs are processed in Australian data centres, or on your own hardware.

No training on your data

Your data is never used to train or improve our models — or anyone else’s.

No retention by default

Requests are processed and discarded, not logged for review or analytics.

No lock-in

Open models you can take with you, behind an API you already know.

Models

Open models. No vendor lock-in.

Frontier-level open models through one platform — 300+ behind one API, including your own private fine-tunes.
deepseek-v4-flash
deepseek-v4-flash
Agents, reasoning and coding
kimi-k3
kimi-k3
Reads whole document sets — 1M context
qwen-3.8
qwen-3.8
Chat, RAG and everyday workloads
gemma-4
gemma-4
High volume at the lowest cost
One API · any model
Your app
POST /v1/chat
Xinference
one API
AU or self-hosted
deepseek-v4-flash deepseek-v4-flash
kimi-k3 kimi-k3
qwen-3.8 qwen-3.8
gemma-4 gemma-4
+ 300+ model options
The economics

Same workload. A third of the bill.

Open models on dedicated infrastructure cost a fraction of closed-source API pricing — and the cost per token is predictable, not metered against someone else’s margins.

*Indexed to closed-source API = 100. Customer-reported infrastructure comparisons, illustrative at the top of the reported range.

Closed-source API100
Hosted open models55
Private models via Xinference30
Up to 70% lower inference cost*
Solutions

What you can build on Xinference.

From experimentation to production — every model type on one platform, behind one API.

LLMs

Chat, code and reasoning with long-context streaming.

Embeddings & search

High-throughput dense retrieval for RAG, semantic search and recommendations.

Transcription (STT)

Real-time streaming speech-to-text.

Text-to-speech

Low-latency voices for phone calls and voice agents.

Image generation

FLUX, SDXL and SD3 with fast diffusion.

Multi-model pipelines

Embedding, rerank and LLM on one runtime.

Deployment options

Deployed in our Australian cloud or your own.

01

Private Australian cloud

Your own dedicated environment in Australian data centres. Nothing shared, nothing offshore.

✓Dedicated GPUs, not a shared pool
✓Running within days
02

Your own infrastructure

Install into your data centre or your cloud account, behind your own perimeter.

✓Your VPC or your racks, your keys
✓Built for regulated workloads
03

AWS Marketplace

Deploy through AWS Marketplace and keep billing on the AWS account you already have.

✓Counts toward AWS spend commitments
✓Procurement already done

Sovereign AI starts here.

Coming from OpenAI, Claude or Bedrock — or starting from nothing? We will show you what your workload looks like on private models in Australia, and what it costs.

Xinference Australia Pty Ltd
Models Deployment Privacy Terms