Talk to an engineer
Solutions

One runtime for every modality.

LLMs, embeddings, speech and image generation, served from the same control plane, the same API and the same GPUs.

LLMs

Chat, code, reasoning and tool calls, with streaming and long context.

Pre-tuned models
gpt-oss 120BDeepSeek V4Llama 3.3 70BQwen 3.5Mistral LargeKimi K2-Instruct
Learn more

Embeddings

High-throughput dense retrieval and reranking, batched and cached.

Pre-tuned models
BGE-large-enJina Embeddings v3NV-Embed v2BGE RerankerCohere EmbedE5-Mistral-7B
Learn more

Transcription

Real-time speech-to-text with diarization, tuned for voice agents.

Pre-tuned models
Whisper Large v3Distil-WhisperCanary-1BParakeet TDTSeamlessM4T v2
Learn more

Text-to-Speech

Low-latency voice synthesis for conversational agents and IVR.

Pre-tuned models
KokoroXTTS v2StyleTTS2BarkOpenVoice v2
Learn more

Image Generation

FLUX, SDXL and SD3, served as one endpoint with hot-swappable LoRAs.

Pre-tuned models
FLUX.1 DevFLUX.1 SchnellSDXLSD3 MediumPlayground v3
Learn more

Multi-model Pipelines

Embedding, rerank and LLM chained on one runtime, sharing GPUs and metrics.

Pre-tuned models
Router · Mistral 7BClassifier · BGEReasoning · QwQ-32BEval judge · Llama 8BGuardrail · Mistral
Learn more
One runtime

Every modality, one control plane.

Xinference exposes LLMs, embeddings, speech and image models through one API, scheduled across one pool of GPUs.

  • Same OpenAI-compatible API. Only the model name changes.
  • Same GPU pool, scheduled across model types by one engine.
  • Same observability: TTFT, TPOT, throughput and cost.
See the developer docs
Your app
One API call
Xinference runtime
Same endpoint, same GPUs
Any modality

One runtime.
Every modality.

From a 7B embedding model to a 120B reasoning chain, same CLI, same SDK, same GPUs.