What Do You Need to Run GLM Locally? VRAM by Version
GLM currently has 5 versions you can deploy locally. The smallest, GLM 4.7 Flash, needs 4GB of VRAM at Q4; the largest, GLM 5.3, needs 41GB. The cheapest device that gets one running is the Tesla V100 16GB, at about $131.
VRAM needed for each GLM version
Xianyu
| Version | Released | Context | |||
|---|---|---|---|---|---|
| GLM 4.7 FlashMoE | 30B (3B active) | Jan 19, 2026 | 198K | 4plus 17 GB of RAM | Tesla V100 16GB$131 · 16 GB · Experts offloaded |
| GLM 4.5 AirMoE | 106B (12B active) | Jul 20, 2025 | 128K | 7plus 56 GB of RAM | Tesla V100 16GB$131 · 16 GB · Experts offloaded |
| GLM 5.3 FlashMoE | 320B (18B active) | Aug 25, 2026 | 1M | 30plus 174 GB of RAM | Instinct MI50 32GB$405 · 32 GB · Experts offloaded |
| GLM 5.2MoE | 744B (40B active) | Jun 16, 2026 | 1M | 41plus 406 GB of RAM | CMP 170HX 64GB$2,089 · 64 GB · Experts offloaded |
| GLM 5.3MoE | 744B (40B active) | Aug 25, 2026 | 1M | 41plus 406 GB of RAM | CMP 170HX 64GB$2,089 · 64 GB · Experts offloaded |
Minimum VRAM = Q4_K_M weights + KV cache for an 8K context + 1GB runtime overhead, rounded up. The cheapest device is the first one that fits when sorting by median price on the current price basis.
About GLM
GLM is published by Z.ai / Zhipu AI under MIT / GLM-5.3 License. This page covers 5 versions ranging from 30B to 744B parameters. New official releases are scanned weekly; architecture parameters and sourced or estimated weight sizes are reviewed before publication.
VendorZ.ai / Zhipu AI
LicenseMIT / GLM-5.3 License
Official sitehttps://z.ai/
Hugging Face orgzai-org