LLMs
Chat, code, reasoning and tool calls, with streaming and long context.
LLMs, embeddings, speech and image generation, served from the same control plane, the same API and the same GPUs.
Chat, code, reasoning and tool calls, with streaming and long context.
High-throughput dense retrieval and reranking, batched and cached.
Real-time speech-to-text with diarization, tuned for voice agents.
Low-latency voice synthesis for conversational agents and IVR.
FLUX, SDXL and SD3, served as one endpoint with hot-swappable LoRAs.
Embedding, rerank and LLM chained on one runtime, sharing GPUs and metrics.
Xinference exposes LLMs, embeddings, speech and image models through one API, scheduled across one pool of GPUs.
Deploy any model to your own GPUs with one command and enterprise governance built in.
Explore the product →Paged KV-cache, continuous batching and autoscaling, tuned for latency.
See the platform →SDKs, one-line migration and a model catalog ready to launch.
Read the docs →From a 7B embedding model to a 120B reasoning chain, same CLI, same SDK, same GPUs.