Skip to content

Sky Forge · cost calculator

Is cloud GPU rental worth it vs the box under your desk?

Three ways to run an open LLM: buy the hardware and pay for power, rent a GPU by the second, or pay a frontier API per token. We put the amortization, the watts, and the published rates on the table from sourced numbers. The math is here — you decide.

sourced estimates · as of 2026-07-20 · no marketing math

01The calculator

Pick a model, a rig, and your usage.

Outputs are sourced estimates, not quotes. Every third-party figure carries its source and as-of date; only the two Skyforge rates are ours.

Home rig to self-host on

Model to run

Frontier API reference

Usage

2.0M tokens/day

50K100.0M
Way to run itWhat runs$ / monthBreak-even
Self-hostRTX 5090 (32 GB)
Qwen3 32B (Q4 quant)32B dense, 4-bit quant
$51.58lowest $/mo
$35.00 amort + $16.58 power · estimate
Rent on Skyforge1× RTX PRO 6000 (96 GB)*
DeepSeek-R1-Distill 32Bnearest Skyforge catalog model (same 32B dense class), served at full precision (bf16) rather than 4-bit quant.
$224
$2.25/GPU-hr · active hours only · estimate
month 7Launch →
Frontier APIAnthropic
Claude Sonnet 5closed frontier model (different weights)
$365
$3.00/$15.00 per 1M in/out · estimate
month 4
Estimates · prices as of 2026-07-20 · assumptions editable below~4561 tok/s decode (midpoint used), community benchmarks* Skyforge doesn't rent 5090s — this is the nearest rentable card, billed as if it ran no faster than your 5090 (it's faster).

02The ceiling

The frontier you can't run at home.

The biggest home rig here tops out around gpt-oss-120b-class weights. These open models are heavier than that by an order of magnitude — renting datacenter GPUs is the only practical way to run the full open frontier weights yourself.

01

DeepSeek V4 Pro

1.6T MoE · 49B act · 16× 141 GBThe largest model in the catalog, for the hardest problems. Its minimum shape is 16× 141 GB of datacenter GPUs — beyond any single box under a desk.Rent DeepSeek V4 Pro on Skyforge →

02

GLM 5.2

744B MoE · 8× 141 GBHeaviest coding and agentic work, with a 1M-token context. Its minimum shape is 8× 141 GB of datacenter GPUs — beyond any single box under a desk.Rent GLM 5.2 on Skyforge →

03

Mistral Large 3

675B MoE · 41B act · 8× 141 GBMistral's frontier model — bring your Hugging Face token. Its minimum shape is 8× 141 GB of datacenter GPUs — beyond any single box under a desk.Rent Mistral Large 3 on Skyforge →

04

Kimi K2.7-Code

1T MoE · 32B act · 8× 80 GBMoonshot's code-specialized frontier model. Its minimum shape is 8× 80 GB of datacenter GPUs — beyond any single box under a desk.Rent Kimi K2.7-Code on Skyforge →

On Kimi: Kimi K3 has open weights announced for July 27, 2026. Moonshot platform pricing · as of 2026-07-20.

03Methodology

Every formula, and the constants it runs on.

The disclosures below render the same constants the calculator computes with — so what you read here can't drift from the math. Expand each to see the formula and its inputs, with sources.

Hardware, amortized

monthly = (price − price × residual%) ÷ 36 months

We spread the buy price over a 36-month horizon, net of what the card is worth at resale. Residual value assumed: 80% at 12mo, 70% at 24mo, 60% at 36mo.

Electricity

monthly = (load W × active-hrs + idle W × idle-hrs) ÷ 1000 × $0.188/kWh

Load watts apply while the rig is decoding; idle watts the rest of the month. Rate used: $0.188/kWh.U.S. EIA Electric Power Monthly, April 2026 average residential rate (18.83¢/kWh) · as of 2026-07-20

Skyforge rental

monthly = active-hrs × 1.25 × published rate

Active hours are inflated by a ×1.25 overhead factor for real-world request padding and non-decode time. That overhead is applied to the Skyforge line only — a deliberately conservative assumption against ourselves. The two rates are the published on-demand prices, mirroring /pricing:

  • 1× RTX PRO 6000 (96 GB) — $2.25/GPU-hr
  • 1× RTX PRO 6000 (96 GB) — $2.25/GPU-hr
  • DGX Spark (1 unit, 128 GB) — $0.59/instance-hr
Frontier API, per token

blended $/1M = input $ × 0.75 + output $ × 0.25

API cost blends input and output rates at a 75/25 input:output token mix. Rates below are published list prices; each links to its source.

Provider · modelIn $/1MOut $/1MBlended $/1M
Z.ai GLM-5.2 · open weights$1.40$4.40$2.15
Alibaba Qwen3.7-Max$2.50$7.50$3.75
Moonshot Kimi K3 · open weights$3.00$15.00$6.00
Anthropic Claude Opus 4.8$5.00$25.00$10.00
Anthropic Claude Sonnet 5$3.00$15.00$6.00
OpenAI GPT-5.6 terra$2.50$15.00$5.63
OpenAI GPT-5.6 sol$5.00$30.00$11.25

List rates as of 2026-07-20 — sourced estimates, not quotes.

04FAQ

Honest questions about the numbers.

05Get started

Rent it by the second, or check the rate card.

On-demand is self-serve from the console — one private OpenAI-compatible endpoint, one key, billed per second. No minimums, terminate anytime.