How is the Model API billed?
Per token, per model. Contact sales for current rates and volume pricing.
Can we move to dedicated compute later?
Yes. Every option runs the same OpenAI-compatible API, so a move is a base URL change, not a rewrite.
Is our data used for training?
No. No training on your data, no retention by default, processed in Singapore data centres or on your own hardware.
Which models can we run?
300+ open models including GLM, DeepSeek, Gemma, Qwen and Llama, plus your own fine-tunes.
Do you support function calling and structured output?
Yes. Function calling, structured output (JSON mode), streaming, embeddings and vision models are all served through the same API.
What are the rate limits?
Default limits apply per model. Higher concurrency is available on request through sales.
Is there an SLA?
Yes. Paid plans come with an uptime SLA and defined support response times. Ask sales for the current terms.