Sky Forge · cost calculator
Is cloud GPU rental worth it vs the box under your desk?
Three ways to run an open LLM: buy the hardware and pay for power, rent a GPU by the second, or pay a frontier API per token. We put the amortization, the watts, and the published rates on the table from sourced numbers. The math is here — you decide.
sourced estimates · as of 2026-07-20 · no marketing math
01The calculator
Pick a model, a rig, and your usage.
Outputs are sourced estimates, not quotes. Every third-party figure carries its source and as-of date; only the two Skyforge rates are ours.
Model to run
Frontier API reference
Usage
2.0M tokens/day
| Way to run it | What runs | $ / month | Break-even |
|---|---|---|---|
Self-hostRTX 5090 (32 GB) | Qwen3 32B (Q4 quant)32B dense, 4-bit quant | $51.58lowest $/mo $35.00 amort + $16.58 power · estimate | — |
Rent on Skyforge1× RTX PRO 6000 (96 GB)* | DeepSeek-R1-Distill 32Bnearest Skyforge catalog model (same 32B dense class), served at full precision (bf16) rather than 4-bit quant. | $224 $2.25/GPU-hr · active hours only · estimate | month 7Launch → |
Frontier APIAnthropic | Claude Sonnet 5closed frontier model (different weights) | $365 $3.00/$15.00 per 1M in/out · estimate | month 4 |
02The ceiling
The frontier you can't run at home.
The biggest home rig here tops out around gpt-oss-120b-class weights. These open models are heavier than that by an order of magnitude — renting datacenter GPUs is the only practical way to run the full open frontier weights yourself.
DeepSeek V4 Pro
1.6T MoE · 49B act · 16× 141 GBThe largest model in the catalog, for the hardest problems. Its minimum shape is 16× 141 GB of datacenter GPUs — beyond any single box under a desk.Rent DeepSeek V4 Pro on Skyforge →
GLM 5.2
744B MoE · 8× 141 GBHeaviest coding and agentic work, with a 1M-token context. Its minimum shape is 8× 141 GB of datacenter GPUs — beyond any single box under a desk.Rent GLM 5.2 on Skyforge →
Mistral Large 3
675B MoE · 41B act · 8× 141 GBMistral's frontier model — bring your Hugging Face token. Its minimum shape is 8× 141 GB of datacenter GPUs — beyond any single box under a desk.Rent Mistral Large 3 on Skyforge →
Kimi K2.7-Code
1T MoE · 32B act · 8× 80 GBMoonshot's code-specialized frontier model. Its minimum shape is 8× 80 GB of datacenter GPUs — beyond any single box under a desk.Rent Kimi K2.7-Code on Skyforge →
On Kimi: Kimi K3 has open weights announced for July 27, 2026. Moonshot platform pricing · as of 2026-07-20.
03Methodology
Every formula, and the constants it runs on.
The disclosures below render the same constants the calculator computes with — so what you read here can't drift from the math. Expand each to see the formula and its inputs, with sources.
Hardware, amortized
monthly = (price − price × residual%) ÷ 36 months
We spread the buy price over a 36-month horizon, net of what the card is worth at resale. Residual value assumed: 80% at 12mo, 70% at 24mo, 60% at 36mo.
- RTX 5090 (32 GB) — $2,900–$3,400Best Value GPU price tracker, July 2026 (MSRP $1,999 is effectively unobtainable) · as of 2026-07-20
- RTX PRO 6000 Blackwell (96 GB) — $11,360–$14,499NVIDIA list $13,250; PNY/B&H street range, July 2026 · as of 2026-07-20
- DGX Spark (128 GB unified) — $4,699NVIDIA store price, July 2026 (raised from $3,999 launch) · as of 2026-07-20
Electricity
monthly = (load W × active-hrs + idle W × idle-hrs) ÷ 1000 × $0.188/kWh
Load watts apply while the rig is decoding; idle watts the rest of the month. Rate used: $0.188/kWh.U.S. EIA Electric Power Monthly, April 2026 average residential rate (18.83¢/kWh) · as of 2026-07-20
- RTX 5090 (32 GB) — 575 W load / 65 W idleNVIDIA spec TDP · as of 2026-07-20
- RTX PRO 6000 Blackwell (96 GB) — 600 W load / 65 W idleNVIDIA spec TDP · as of 2026-07-20
- DGX Spark (128 GB unified) — 100 W load / 10 W idlemeasured during LLM inference, ~60–100 W; 240 W system max · as of 2026-07-20
Skyforge rental
monthly = active-hrs × 1.25 × published rate
Active hours are inflated by a ×1.25 overhead factor for real-world request padding and non-decode time. That overhead is applied to the Skyforge line only — a deliberately conservative assumption against ourselves. The two rates are the published on-demand prices, mirroring /pricing:
- 1× RTX PRO 6000 (96 GB) — $2.25/GPU-hr
- 1× RTX PRO 6000 (96 GB) — $2.25/GPU-hr
- DGX Spark (1 unit, 128 GB) — $0.59/instance-hr
Frontier API, per token
blended $/1M = input $ × 0.75 + output $ × 0.25
API cost blends input and output rates at a 75/25 input:output token mix. Rates below are published list prices; each links to its source.
| Provider · model | In $/1M | Out $/1M | Blended $/1M |
|---|---|---|---|
| Z.ai GLM-5.2 · open weights | $1.40 | $4.40 | $2.15 |
| Alibaba Qwen3.7-Max | $2.50 | $7.50 | $3.75 |
| Moonshot Kimi K3 · open weights | $3.00 | $15.00 | $6.00 |
| Anthropic Claude Opus 4.8 | $5.00 | $25.00 | $10.00 |
| Anthropic Claude Sonnet 5 | $3.00 | $15.00 | $6.00 |
| OpenAI GPT-5.6 terra | $2.50 | $15.00 | $5.63 |
| OpenAI GPT-5.6 sol | $5.00 | $30.00 | $11.25 |
List rates as of 2026-07-20 — sourced estimates, not quotes.
04FAQ
Honest questions about the numbers.
05Get started
Rent it by the second, or check the rate card.
On-demand is self-serve from the console — one private OpenAI-compatible endpoint, one key, billed per second. No minimums, terminate anytime.