New: Model API. Frontier open models, hosted in Australia →

One control plane,
from app to GPU.

One integration that keeps working as models, hardware and traffic change. Deploy, route and govern open models across cloud and on-prem infrastructure, with one API for your applications.

How it fits

What you get

Model repository

300+ open models, embedding and multimodal, plus your own fine-tunes, deployed from one catalogue.

Routing

Our proprietary smart routing engine optimises cost and performance. We can also pair you with an engineer to tune it for your workload.

Deployment options

Our cloud, your cloud, or your data centre, including fully offline environments.

Governance

Role-based access at organisation and resource level, audit logs, SSO via Google Workspace, and live monitoring of TTFT and TPOT.

Engines under the hood: vLLM · SGLang · MLX · Transformers · one OpenAI-compatible REST API with streaming.

Running 50+ AI use cases in production for Yum! Brands.

Where it runs

Xinference Cloud

Your own dedicated environment in Australian data centres.

Your infrastructure

Install into your data centre or your cloud account, behind your own perimeter. Built for regulated workloads.

AWS Marketplace

Deploy through AWS Marketplace and keep billing on the AWS account you already have.

Every model type

LLMs: chat, code and reasoning with long-context streaming.

Embeddings and search: dense retrieval for RAG and semantic search.

Multi-model pipelines: embedding, rerank and LLM on one runtime.

Transcription: real-time streaming speech to text.

Text to speech: low-latency voices for calls and agents.

Image generation: FLUX, SDXL and SD3.

See it on your workload.