Skip to content

Model catalogQwen

Qwen3.8-Flash-Next

Qwen4-architecture preview — ultra-sparse, fast, and multimodal. Served with vLLM behind your own private, OpenAI-compatible /v1 endpoint — billed by the second, terminate anytime.

125B MoE · 6B act · FP8 · 256K context · 4× 96 GB · per-second billing

01Specifications

The specs the launcher enforces.

Architecture, served precision, max context, the minimum GPU shape, and the exact pinned weights — the facts that decide how Qwen3.8-Flash-Next runs and what it costs.

SpecValue
Architecture125B MoE · 6B act
Served precisionFP8
Max context256K
Minimum GPU shape4× 96 GB
Weights servedQwen/Qwen3.8-Flash-Next-FP8@970c569

Runs on 4× 96 GB · billed per second — final cost depends on the GPU shape.

Full pricing →

02How it runs

A private endpoint, not a shared API.

Launching Qwen3.8-Flash-Next runs the model on your own instance and gives you a private, OpenAI-compatible /v1 endpoint and API key. Point any OpenAI-compatible SDK at your base URL, change the model name, and existing code works. The endpoint is yours alone — terminate the instance and it goes away with it.

Billing is by the second on the 4× 96 GB shape the launcher enforces, from boot to terminate. No commitment and no tier gates.

03Get started

Launch Qwen3.8-Flash-Next now.

One private OpenAI-compatible endpoint, one key, billed by the second. The console opens with Qwen3.8-Flash-Next preselected.