Mid-size company build

Running GLM-5.2

Z.ai (Zhipu AI) · MIT · 744B parameters

The memory math: 744B parameters × 2 GB/B = 1,488 GB baseline, + 20% overhead for KV cache/context = 1,786 GB minimum required.
Recommended build

1x NVIDIA GB200 NVL72

A full rack-scale cluster instead of a pair of servers: over 7x the memory headroom this model actually needs, so the same rack can also serve many more simultaneous users, run a second model alongside it, or absorb a hardware failure without downtime.

Combined GPU memory
13,824 GB
Total price
$3,000,000
Combined power draw
120.0 kW
Monthly electricity
$16,243

1 × NVIDIA GB200 NVL72 at 120,000 W each = 120,000 W total. Run 24/7, that's 2,880.0 kWh/day — about the same continuous draw as 99.3 average U.S. homes (~29 kWh/day/home, EIA). Monthly cost assumes 18.8¢/kWh, the 2026 U.S. residential average.

See full specs for NVIDIA GB200 NVL72 →

Request this build

No payment, no account — just tell us your name and we'll follow up with a real quote.

Building for a small startup instead? See the startup build for GLM-5.2 →