Sovereign inference. Run any model.

Run AI models on your infrastructure. One platform to deploy, optimise and govern open models.

Xinference

An enterprise platform for running open models on infrastructure you control. Run in Australia. Up to 70% cheaper than closed-source APIs.*

300+ open models Latest open models, fast APIs or dedicated GPUs

Amazon Bedrock

AWS's serverless service built for foundation models, billed per token or per hour.

Serverless Per-token or hourly Frontier models and some open models

Where Xinference wins

01

Better economics

A proprietary stack we built ourselves: up to 70% lower cost, 2-4x faster responses.*

02

Hands-on support

We run the proof of concept with you, then provide dedicated support in production.

03

Any infrastructure

Deploy on your own infrastructure or ours. No vendor lock-in provides cost and data control.

04

Access the latest open models

300+ open models: LLMs, embeddings, speech and image, plus your own fine-tunes.

Feature comparison

Last verified: 31 August 2026

What matters Xinference Amazon Bedrock
Model choice 300+ open models Hundreds, varies by region
Lock-in Model, infrastructure and hardware agnostic Tied to AWS infrastructure
Where it runs Your own infrastructure (existing cloud, self-hosted) or ours AWS only
Pricing model Per token, or per GPU/hourPay only for what you use. Custom enterprise rates available. Per token, or hourly reserved
Cost Up to 70% lower* Standard AWS rates
New open models Same API, no re-integration Generally several months behind
Performance 2-4x faster responses* Varies by model
Support PoC, onboarding, dedicated support Separate paid support plan

*Xinference-reported figures, compared with closed-source APIs. Results vary by workload.

Which one fits your team?

Xinference is the better fit if

  • You want to fully own your environment: your cloud, or self-hosted.
  • You want the latest open models as they release.
  • You want engineers running the PoC and rollout with you.

Bedrock is the better fit if

  • You're fully committed to AWS: one account, one bill.
  • You need direct access to frontier models.

Prefer not to manage infrastructure at all? Our hosted model APIs cover that too.

Common questions

Will our existing code work with Xinference?

Yes. The API is OpenAI-compatible, so for most applications the switch is one line of configuration.

Where does Xinference run?

On your own infrastructure (your existing cloud, or self-hosted) or on ours, run in Australia. You choose where it runs. It's also available through AWS Marketplace.

What does no lock-in mean in practice?

Your applications talk to one OpenAI-compatible API. You can change models, serving engines or GPUs without rebuilding your applications.

How is Xinference priced?

Per token for hosted model APIs, or per GPU/hour for dedicated deployments: you pay only for what you use. Custom enterprise rates are available for larger workloads.

Sovereign AI starts here.

Coming from OpenAI, Claude or Bedrock, or starting from nothing? We will show you what your workload looks like on private models and what it costs.