Model catalogKimi
Kimi K3
Moonshot's flagship agentic model, with a 1M-token context. Served with vLLM behind your own private, OpenAI-compatible /v1 endpoint — billed by the second, terminate anytime.
2.8T MoE · 104B act · MXFP4 · 1M context · 16× 141 GB · per-second billing
01Specifications
The specs the launcher enforces.
Architecture, served precision, max context, and the minimum GPU shape — the numbers that decide how Kimi K3 runs and what it costs.
| Spec | Value |
|---|---|
| Architecture | 2.8T MoE · 104B act |
| Served precision | MXFP4 |
| Max context | 1M |
| Minimum GPU shape | 16× 141 GB |
Runs on 16× 141 GB · billed per second — final cost depends on the GPU shape.
Full pricing →02How it runs
A private endpoint, not a shared API.
Launching Kimi K3 runs the model on your own instance and gives you a private, OpenAI-compatible /v1 endpoint and API key. Point any OpenAI-compatible SDK at your base URL, change the model name, and existing code works. The endpoint is yours alone — terminate the instance and it goes away with it.
Billing is by the second on the 16× 141 GB shape the launcher enforces, from boot to terminate. No commitment and no tier gates.
03Get started
Launch Kimi K3 now.
One private OpenAI-compatible endpoint, one key, billed by the second. The console opens with Kimi K3 preselected.