SkyForge · model catalog
Every model we serve, specified.
Eighteen curated open-weights models — GLM, Qwen, DeepSeek, Llama 4, MiniMax, Kimi, Gemma, Mistral — each pinned to a versioned template and served behind your own private /v1 endpoint (vLLM for text, vLLM-Omni for video). The specs below are the ones the launcher actually enforces: architecture, served precision, max context, and the minimum GPU shape.
19 curated models · served with vLLM · one key, every endpoint · per-second billing
01Launch it
Self-serve, on-demand.
Everything in the catalog launches on an on-demand GPU shape from the live rate card — single-GPU to multi-GPU nodes. Pick a model below; the launcher puts it on the right shape.
On-demand · per-second
Single-GPU instances for inference, fine-tuning, and dev work. Launch an open-weights model from a template and pay by the second.
02Models
Pick what you're building.
Choose a category to see the models tuned for it — architecture, served precision, max context, and the minimum GPU shape the launcher enforces. Every row launches in one click.
| Model | Architecture | Served | Max context | Min GPUs | |
|---|---|---|---|---|---|
| 9B dense | bf16 | 128K | 1× 24 GB | ||
| 27B dense | bf16 | 256K | 1× 80 GB | ||
| 27B dense | bf16 | 256K | 1× 80 GB |
Served with vLLM behind your own private, OpenAI-compatible /v1 endpoint.
03Bring your own
Your code, your weights, same meter.
The GPU Dev / JupyterLab template is a notebook workspace with PyTorch and CUDA — not an inference endpoint. Bring your own code, weights, and experiments instead of serving a catalog model, billed by the second like everything else.
04Open-weights families
The families we serve.
Every model in the catalog ships open weights and runs on the same OpenAI-compatible serving path. Pick the family you trust; run it without sending your prompts to a closed frontier lab.
GLM
GLM-4 9B · GLM 5.3 Flash · GLM 5.2Zhipu's GLM line — from a light single-GPU chat model to 320B and 744B mixture-of-experts models with 1M-token contexts for coding and agents.
Qwen
Qwen3.8 27B · Qwen3.6 27B · 35B-A3B · Qwen3.8-Flash-Next · Qwen3.5 397B-A17BAlibaba's Qwen family, spanning a single-GPU dense model, an efficient MoE, a Qwen4-architecture Flash preview, and a large frontier MoE for demanding analysis.
DeepSeek
R1-Distill 32B · V4 Flash · V4 ProDeepSeek reasoning and mixture-of-experts models, from affordable single-GPU R1 distillation up to the largest model in the catalog.
Llama
Llama 4 ScoutMeta's Llama 4 Scout, a mixture-of-experts general assistant. Gated on Hugging Face — launch with your own token.
MiniMax
MiniMax M3 · MiniMax H3A mixture-of-experts model built for very long-context reasoning, with a window up to 1M tokens — plus H3, the omni-modal video + audio generation model.
Kimi
Kimi K2.7-CodeMoonshot's code-specialized frontier model for the heaviest coding and agentic work.
Gemma
Gemma 3 27BGoogle's Gemma 3 Instruct. Gated on Hugging Face — launch with your own token.
Mistral
Mistral Large 3Mistral's frontier mixture-of-experts model. Gated on Hugging Face — launch with your own token.
05FAQ
Questions, answered.
06Get started
Launch a model, or tell us what you need next.
Every model above launches from the console today — one private OpenAI-compatible endpoint and one key, billed by the second. Want a model or family we don't list yet? Get on the list and tell us what you're building.
Eyeing the new desk-side hardware — Mac Studio M5 Ultra, DGX Spark, DGX Station? Join the hardware waitlist →