128GB unified memory
NVIDIA DGX Spark
GB10 Grace Blackwell on the desk — CUDA on 128GB of unified memory.
Pilot allocation, first-come
Qwen3.8 27B
27B dense at bf16 (~54GB) fits with headroom for a 256K-token KV cache.
DeepSeek-R1-Distill 32B
32B dense at bf16 (~64GB) — step-by-step reasoning on a single unit.
Gemma 3 27B
27B dense at bf16 (~54GB) runs on the 128GB pool with room to spare.