Skip to content

SkyForge · open-weights model hosting

Your models. Your endpoint.
Your GPU.

Launch any open-weights model — GLM, Qwen, DeepSeek, Llama 4, Mistral — on a dedicated GPU with a private, OpenAI-compatible /v1 endpoint and one API key. Billed per second. Terminate anytime.

OpenAI-compatible /v1Per-second billingOne key, every instance19 curated models + bring-your-own
your endpoint — private /v1
# launched from the console · boots in minutes
$ curl https://swift-mesa-4821.skyforgecompute.com/v1/chat/completions \
    -H "Authorization: Bearer $SKYFORGE_API_KEY" \
    -d '{"model":"glm-4-9b","messages":[…]}'

{"choices":[{"message":{"content":"Hello from your own GPU."}}]}

01Models

The open-weights frontier, as a spec sheet.

Curated, versioned templates — from a 9B chat model on a single 24 GB GPU to 2.8T-parameter frontier MoE on an 8× 288 GB node. Each launches to its own private endpoint.

ModelArchitectureServedMax contextMin GPUs
GLM-4 9B
tier Ssingle GPU
9B densebf16128K1× 24 GBLaunch →
Qwen3.6 27B
27B densebf16256K1× 80 GBLaunch →
Qwen3.8 27B
27B densebf16256K1× 80 GBLaunch →
Gemma 3 27Bgated
27B densebf16128K1× 80 GBLaunch →
DeepSeek-R1-Distill 32B
32B densebf16128K1× 80 GBLaunch →
Qwen3.6 35B-A3B
35B MoE · 3B actbf16256K1× 96 GBLaunch →
Llama 4 Scoutgated
tier Mmulti-GPU
109B MoE · 17B actFP8256K2× 80 GBLaunch →
DeepSeek V4 Flash
284B MoE · 13B actFP8256K4× 96 GBLaunch →
Qwen3.8-Flash-Next
125B MoE · 6B actFP8256K4× 96 GBLaunch →
Qwen3.5 397B-A17B
397B MoE · 17B actFP8256K8× 80 GBLaunch →
MiniMax M3
428B MoE · 23B actFP81M8× 80 GBLaunch →
GLM 5.3 Flash
320B MoE · 18B actFP81M8× 80 GBLaunch →
DeepSeek V4.1 Flash
552B MoE · 8–16B actMXFP41M8× 96 GBLaunch →
MiniMax H3
33B dense omni · videobf164× 141 GBLaunch →
Kimi K2.7-Code
tier Ffrontier cluster
1T MoE · 32B actINT4256K8× 80 GBLaunch →
Mistral Large 3gatedlimited supply
675B MoE · 41B actFP8256K8× 141 GBLaunch →
GLM 5.2limited supply
744B MoEFP81M8× 141 GBLaunch →
DeepSeek V4 Prolimited supply
1.6T MoE · 49B actFP81M8× 288 GBLaunch →
Kimi K3limited supply
2.8T MoE · 104B actMXFP41M8× 288 GBLaunch →

Text models served with vLLM behind a private, OpenAI-compatible /v1 endpoint; MiniMax H3 serves its video API (/v1/videos) via vLLM-Omni. Plus JupyterLab + PyTorch for your own code.

Full catalog →

New — hardware pilot

Mac Studio M5 Ultra, DGX Spark, DGX Station — be first in line.

Desk-side machines with the unified memory to run frontier open-weights models locally. Allocations are first-come.

Join the waitlist

02How it works

Intent to a live endpoint in three steps.

STEP 1

Pick a model by intent

Chat, coding, reasoning, or your own notebook — the guided flow surfaces the open-weights model that fits, with a plain-English why.

STEP 2

Set a budget tier

Value, Recommended, or Performance. We map it to the right GPU shape and show the per-second rate before you commit.

STEP 3

Call your endpoint

A private, OpenAI-compatible /v1 with one API key across every instance. Point the SDK you already use at it — no rewrites.

base_url = https://<your-slug>.skyforgecompute.com/v1 · api_key = one SkyForge key, every instance

03Pricing

Pay by the second. Stop, and you stop paying.

No minimums, no tier gates, no reserved commitment to start. Live accrued, burn-rate, and projected spend in the console.

The catalog

every GPU type

$0.28/ GPU-hr

From-rates across 20 GPU types — workstation cards to datacenter Blackwell, platform fee included, billed per second.

RTX PRO 6000

inference workhorse

$2.85/ GPU-hr

96 GB GDDR7 per GPU, 1–8 GPUs per node. Launch any Tier-S template on a single GPU — billed per second from ready to terminate.

04Why SkyForge

A private endpoint you actually control.

01

One operator, not a marketplace

Your job runs on our own fleet, start to finish. It won't vanish mid-run on an unverified third-party host.

02

OpenAI-compatible, zero rewrites

Every endpoint speaks the OpenAI /v1 API. Switch a base URL and ship — same SDK, same tooling, your weights.

03

Per-second, no lock-in

Pay by the second and terminate anytime. No minimums, no reserved commitment to get started, live spend in the console.

04

Open weights you control

GLM, Qwen, DeepSeek, Llama 4, Mistral, Kimi — the model runs on your instance and the key is yours.

05 — Enterprise

Outgrown on-demand? We'll build with you.

  • Reserved & committed capacity for sustained workloads, allocated through our team with a named point of contact.

  • Buy-through-us hardware. Few teams hold Supermicro or vendor accounts — we can procure the servers and stand them up for you. In planning; tell us what you need.

Where this is going

From brokered compute to our own datacenter.

Today we run on brokered capacity we source and operate, so you can launch in minutes without waiting on hardware.

In parallel: real land, secured power, modular container deployments, and an open Supermicro purchase order — we build the datacenter as demand proves in. No stranded capacity, and no claims we can't stand behind.

Read the plan →

06Early access

Get launch news and capacity updates.

Tell us what you're building and we'll follow up — plus a heads-up as reserved capacity and the hardware path come online.

Request early access

No spam. Just launch news and capacity updates.

What's the use case?

Launch your first model in minutes.

Pick a model, set a budget, and get a private OpenAI-compatible endpoint with one key. Per-second billing, terminate anytime.