What Do You Need to Run GLM Locally? VRAM by Version
GLM currently has 5 versions you can deploy locally. The smallest, GLM 4.7 Flash, needs 4GB of VRAM at Q4; the largest, GLM 5.3, needs 41GB. The cheapest device that gets one running is the Radeon RX 7900 XT, with no used-price data yet.
VRAM needed for each GLM version
5 versions| Version | Released | Context | ||||
|---|---|---|---|---|---|---|
| GLM 4.7 FlashMoE | 30B (3B active) | Jan 19, 2026 | 198K | 4plus 17 GB of RAM | Radeon RX 7900 XT— · 20 GB | 1,935,018 |
| GLM 4.5 AirMoE | 106B (12B active) | Jul 20, 2025 | 128K | 7plus 56 GB of RAM | CMP 170HX 64GB (modded)— · 64 GB | 114,264 |
| GLM 5.3 FlashMoE | 320B (18B active) | Aug 25, 2026 | 1M | 30plus 174 GB of RAM | Mac Studio M3 Ultra 512GB— · 512 GB | 654,957 |
| GLM 5.2MoE | 744B (40B active) | Jun 16, 2026 | 1M | 41plus 406 GB of RAM | Mac Studio M3 Ultra 512GB— · 512 GB | 1,134,389 |
| GLM 5.3MoE | 744B (40B active) | Aug 25, 2026 | 1M | 41plus 406 GB of RAM | Mac Studio M3 Ultra 512GB— · 512 GB | 303,534 |
Minimum VRAM = Q4_K_M weights + KV cache for an 8K context + 1GB runtime overhead, rounded up. The cheapest device is the first one that fits when sorting by median used price on the current price basis.
About GLM
GLM is published by Z.ai / Zhipu AI under MIT / GLM-5.3 License. This page covers 5 versions ranging from 30B to 744B parameters; structure parameters and quantized sizes are synced weekly from Hugging Face.
VendorZ.ai / Zhipu AI
LicenseMIT / GLM-5.3 License
Official sitehttps://z.ai/
Hugging Face orgzai-org