Run AI models on your infrastructure. One platform to deploy, optimise and govern open models.
An enterprise platform for running open models on infrastructure you control. Run in Australia. Up to 70% cheaper than closed-source APIs.*
AWS's serverless service built for foundation models, billed per token or per hour.
A proprietary stack we built ourselves: up to 70% lower cost, 2-4x faster responses.*
We run the proof of concept with you, then provide dedicated support in production.
Deploy on your own infrastructure or ours. No vendor lock-in provides cost and data control.
300+ open models: LLMs, embeddings, speech and image, plus your own fine-tunes.
Last verified: 31 August 2026
| What matters | ||
|---|---|---|
| Model choice | ✓300+ open models | ✓Hundreds, varies by region |
| Lock-in | ✓Model, infrastructure and hardware agnostic | ✓Tied to AWS infrastructure |
| Where it runs | ✓Your own infrastructure (existing cloud, self-hosted) or ours | ✓AWS only |
| Pricing model | ✓Per token, or per GPU/hourPay only for what you use. Custom enterprise rates available. | ✓Per token, or hourly reserved |
| Cost | ✓Up to 70% lower* | ✓Standard AWS rates |
| New open models | ✓Same API, no re-integration | ✓Generally several months behind |
| Performance | ✓2-4x faster responses* | ✓Varies by model |
| Support | ✓PoC, onboarding, dedicated support | ✓Separate paid support plan |
*Xinference-reported figures, compared with closed-source APIs. Results vary by workload.
Prefer not to manage infrastructure at all? Our hosted model APIs cover that too.
Yes. The API is OpenAI-compatible, so for most applications the switch is one line of configuration.
On your own infrastructure (your existing cloud, or self-hosted) or on ours, run in Australia. You choose where it runs. It's also available through AWS Marketplace.
Your applications talk to one OpenAI-compatible API. You can change models, serving engines or GPUs without rebuilding your applications.
Per token for hosted model APIs, or per GPU/hour for dedicated deployments: you pay only for what you use. Custom enterprise rates are available for larger workloads.
Coming from OpenAI, Claude or Bedrock, or starting from nothing? We will show you what your workload looks like on private models and what it costs.