A self-hosted alternative to the OpenAI API is an inference server you run on your own infrastructure that accepts the same requests as OpenAI’s API. Your application keeps its code and its OpenAI SDK. You change the base URL and the API key, and requests go to open models running in your environment.
What “OpenAI-compatible” means
The server accepts the same request format as OpenAI’s API, so the official OpenAI SDKs and most tools built on them keep working. Before you switch, check that the server supports the features your app relies on, such as streaming, function calling or structured output.
How to switch in three steps
- Choose open models that fit each task, for example Qwen, DeepSeek, Llama or Gemma.
- Deploy them on your own servers or in your own cloud account.
- Point your application at the new endpoint, and compare outputs on your own prompts before you move traffic.
from openai import OpenAI
client = OpenAI(base_url="https://your-xinference-host/v1", api_key="YOUR_KEY")
What changes and what stays the same
| OpenAI API | Self-hosted, OpenAI-compatible | |
|---|---|---|
| Application code | Your code | The same code, with a new base URL and key |
| Models | OpenAI's models | Open models you choose |
| Where requests run | OpenAI's infrastructure | Your servers or cloud account |
| Cost | Per token | The compute you run |
| Operations | Handled by OpenAI | Your team or an inference platform |
How Xinference fits
Xinference runs in your own environment: your VPC, your data centre or your servers. One OpenAI-compatible API sits in front of more than 300 individual models, and you choose the serving engine for each model: vLLM, SGLang, llama.cpp, Transformers or MLX.
If you are not ready to run your own hardware, you can start with our Model API by the token and move later. Every option runs the same API, so a move is a base URL change, not a rewrite. See Dedicated Inference for running it in your own environment.
Questions
Do the official OpenAI SDKs work?
Yes, as long as the server exposes an OpenAI-compatible API. You change the base URL and the API key.
Will open models give the same answers as OpenAI’s models?
Not always. Results depend on the task and the model. Test on your own prompts before you switch production traffic.
Can we keep using OpenAI for some tasks?
Yes. Because both sides use the same API format, you can send some work to OpenAI and the rest to your own models.