Small startup build
Running GLM-5.2
Z.ai (Zhipu AI) · MIT · 744B parameters
The memory math:
744B parameters × 2 GB/B = 1,488 GB baseline,
+ 20% overhead for KV cache/context = 1,786 GB minimum required.
Recommended build
2x NVIDIA DGX H200
Two pre-integrated 8-GPU servers is the cheapest combination that clears the memory bar for this model. It's still two whole servers — a 744-billion-parameter model simply doesn't fit on anything smaller.
Combined GPU memory
2,256 GB
Total price
$698,000
Combined power draw
20.4 kW
Monthly electricity
$2,761
2 × NVIDIA DGX H200 at 10,200 W each = 20,400 W total. Run 24/7, that's 489.6 kWh/day — about the same continuous draw as 16.9 average U.S. homes (~29 kWh/day/home, EIA). Monthly cost assumes 18.8¢/kWh, the 2026 U.S. residential average.
Why more than one unit? This model needs 1,786 GB
of GPU memory, and a single NVIDIA DGX H200 only has 1,128 GB — so
2 of them are cabled together into what's called a cluster: multiple
servers networked together so they act like one much bigger computer than any single box could be.
See full specs for NVIDIA DGX H200 →
Request this build
No payment, no account — just tell us your name and we'll follow up with a real quote.
Building for a mid-size company instead? See the mid-size build for GLM-5.2 →