We've refreshed the Xinference brand and rebuilt xinference.co from the ground up. This post walks through what changed, and more importantly, why.
A clearer identity: sovereign AI inference
Xinference is the control plane for AI inference. We orchestrate, observe, and scale model deployments for teams that can't compromise on where their data goes. That positioning has always been true of the platform. It's now front and center everywhere you encounter us, from the homepage to the pricing page to this blog.
We're an Australian-owned company, globally trusted, with data residency options in Australia and Singapore. Every deployment path, whether it's the standard bundle or a custom deployment on your own infrastructure, is built around one principle: your data never leaves your environment unless you decide it should.
A new visual system
The new site runs on a rebuilt design system: Open Sans across the board, a single brand blue as the accent color, fully rounded cards and buttons, and no gradients anywhere. It's a deliberately quieter, cleaner look, built to get out of the way of the content: your models, your metrics, your deployment options.
The three-pillar promise that runs through the new site is simple. Sovereign: 100% of inference runs in your environment, and your data never leaves. Cost: up to 70% lower cost versus closed-source APIs. Performance: 2 to 4 times lower latency, up to 4 times more models served per GPU, and no cold starts. Most platforms make you trade one of these for another. Xinference is built to deliver all three at once, on your own infrastructure, through one OpenAI-compatible API.
Clearer, simpler pricing
We've restructured pricing around one simple bundle instead of a confusing tier matrix: a standard bundle at $10K per month, covering unlimited AI agents, 500 users, a dedicated private LLM, and Australian hosting. Additional blocks of 500 users are available as an add-on, and teams that need Xinference running entirely on their own infrastructure, cloud or on-premises, can talk to us about a custom deployment.
Every path routes through the same control plane and the same OpenAI-compatible API. Moving from the standard bundle to a custom deployment as your requirements evolve does not mean re-architecting your application.
Smarter routing, powered by our routing layer
The platform update that underpins the cost story is our routing layer, which we are rolling out to send each request to the cheapest model that clears your quality bar. Combined with KV-cache reuse, continuous batching, and speculative decoding, it's designed to help teams get up to 4 times more useful work out of every GPU, and help open models carry workloads that used to default to a closed-source API by habit rather than necessity.
A new home for the blog
This post lives on the new blog section, which is also new. You'll find engineering write-ups from our founder, tutorials for getting a model running on your own hardware, and updates like this one, all in one place. The goal is straightforward: fewer slide decks changing hands before a technical team gets a straight answer, more of that answer published where anyone evaluating the platform can read it directly.
What hasn't changed
The core commitment is the same as it's always been: no rip-and-replace on day one. Most engagements start as a proof of concept over one to two weeks on the standard bundle, move to a benchmark against your own baseline, and only go to a custom deployment, private cloud, or on-prem once the numbers hold up. The brand looks different. The way we work with teams evaluating the platform does not.

