Guide
What is sovereign inference?
Sovereign inference means running AI models on infrastructure your organisation controls, such as your own cloud account, data centre or servers. Prompts, outputs and model access stay inside your security boundary, and you decide which models run, who can use them and what gets logged. It is the opposite of sending every request to a third-party model API.
Sovereign inference vs a third-party model API
| Third-party model API | Sovereign inference | |
|---|---|---|
| Where requests are processed | The provider’s infrastructure | Your cloud account, data centre or servers |
| Who holds prompts and outputs | The provider, under its own policy | You |
| Model choice | The provider’s catalogue | Any model you deploy, including your own fine-tunes |
| Access control | API keys issued by the provider | Your own roles and single sign-on |
| Cost | Per token, at the provider’s price | The compute you run, planned by your team |
| Operations | Handled by the provider | Handled by your team or an inference platform |
What “sovereign” covers in practice
Sovereignty is not a single setting. It comes down to four questions a security or procurement team will ask:
01
Where does the model run?
02
Which models can we use, and who decides?
03
Who can call the model, and how is that controlled?
04
What is recorded, and who can see it?
If your own team, not the provider, answers each of these, the setup counts as sovereign inference.
Who needs sovereign inference?
Teams usually look at it when one of these is true:
01
The data is sensitive: customer records, health or financial data, legal documents, source code.
02
A contract or internal policy requires data to stay in a given country or environment.
03
Usage is high and steady enough that paying per token costs more than running the compute.
04
The team wants to switch models without rewriting applications or waiting for a provider to add them.
Does sovereign inference mean running everything yourself?
No. The model has to run somewhere you control, but you do not have to build the serving stack. Most teams use an inference platform that handles deployment, scaling and access control, and run it inside their own cloud account or on their own hardware. Less sensitive workloads can still use a hosted API, and an OpenAI-compatible interface lets the two sit side by side.
How Xinference supports sovereign inference
Xinference is an AI inference platform that runs in your own environment: your VPC, your data centre or your servers.
- One OpenAI-compatible API in front of every model. Applications written for the OpenAI API need a new base URL, not a rewrite.
- More than 300 individual models across language, embedding, reranking, speech, image and multimodal tasks, plus your own fine-tunes.
- You choose the serving engine for each model: vLLM, SGLang, llama.cpp, Transformers or MLX.
- Role-based access control, single sign-on through OIDC with Google Workspace, and audit logs of sign-in, access and deployment events. The audit logs do not record prompt or output content.
Questions
Is sovereign inference the same as sovereign AI?
They overlap. Sovereign AI is the broader idea of a country or organisation controlling its AI capability, including data, models and compute. Sovereign inference is the practical part of it: where and how a model runs when it answers a request.
Is sovereign inference the same as on-premises AI?
On-premises is one way to do it. Sovereign inference can also run in your own account with a public cloud provider, as long as you control the environment, the access and the data.
Does sovereign inference cost more?
It depends on volume. At low or uneven usage, paying per token is usually cheaper. At steady, high usage, running your own compute often costs less per request. Model it on your own traffic before you decide.
Do we need to change our applications?
Not if the platform exposes an OpenAI-compatible API. Switching is usually a change to the base URL and API key.
Last updated