Model catalogMistral
Mistral Large 3
Mistral's frontier model — bring your Hugging Face token. Served with vLLM behind your own private, OpenAI-compatible /v1 endpoint — billed by the second, terminate anytime.
675B MoE · 41B act · FP8 · 256K context · 8× 141 GB · per-second billing
01Specifications
The specs the launcher enforces.
Architecture, served precision, max context, the minimum GPU shape, and the exact pinned weights — the facts that decide how Mistral Large 3 runs and what it costs.
| Spec | Value |
|---|---|
| Architecture | 675B MoE · 41B act |
| Served precision | FP8 |
| Max context | 256K |
| Minimum GPU shape | 8× 141 GB |
| Weights served | mistralai/Mistral-Large-3-675B-Instruct-2512@383ffea |
Runs on 8× 141 GB · billed per second — final cost depends on the GPU shape.
Full pricing →02How it runs
A private endpoint, not a shared API.
Launching Mistral Large 3 runs the model on your own instance and gives you a private, OpenAI-compatible /v1 endpoint and API key. Point any OpenAI-compatible SDK at your base URL, change the model name, and existing code works. The endpoint is yours alone — terminate the instance and it goes away with it.
Billing is by the second on the 8× 141 GB shape the launcher enforces, from boot to terminate. No commitment and no tier gates.
Gated model
Mistral Large 3 is a gated repository on Hugging Face. Accept the model's license on Hugging Face, and the launch flow injects your own token so the instance can pull the weights.
Limited GPU supply
8× 141 GB (H200/B200) instances are intermittently out of stock across providers. If the launch flow shows no compute for this shape, check back later — supply returns.
03Get started
Launch Mistral Large 3 now.
One private OpenAI-compatible endpoint, one key, billed by the second. The console opens with Mistral Large 3 preselected.