Skip to content

Model catalogKimi

Kimi K3

Moonshot's flagship agentic model, with a 1M-token context. Served with vLLM behind your own private, OpenAI-compatible /v1 endpoint — billed by the second, terminate anytime.

2.8T MoE · 104B act · MXFP4 · 1M context · 8× 288 GB · per-second billing

01Specifications

The specs the launcher enforces.

Architecture, served precision, max context, the minimum GPU shape, and the exact pinned weights — the facts that decide how Kimi K3 runs and what it costs.

SpecValue
Architecture2.8T MoE · 104B act
Served precisionMXFP4
Max context1M
Minimum GPU shape8× 288 GB
Weights servedmoonshotai/Kimi-K3@9f62e4e

Runs on 8× 288 GB · billed per second — final cost depends on the GPU shape.

Full pricing →

02How it runs

A private endpoint, not a shared API.

Launching Kimi K3 runs the model on your own instance and gives you a private, OpenAI-compatible /v1 endpoint and API key. Point any OpenAI-compatible SDK at your base URL, change the model name, and existing code works. The endpoint is yours alone — terminate the instance and it goes away with it.

Billing is by the second on the 8× 288 GB shape the launcher enforces, from boot to terminate. No commitment and no tier gates.

Limited GPU supply

8× 288 GB (B300) is the only shape with enough total VRAM for this model, and it is the scarcest node on the market — often zero available. If the launch flow shows no compute, check back later.

03Get started

Launch Kimi K3 now.

One private OpenAI-compatible endpoint, one key, billed by the second. The console opens with Kimi K3 preselected.