GPUs tracked41Price history3 dayssince Sep 3, 2026Price points (24h)11,753 pointsBiggest 24h dropBiggest 24h riseData updated Sep 5, 2026

What Do You Need to Run GLM Locally? VRAM by Version

GLM currently has 5 versions you can deploy locally. The smallest, GLM 4.7 Flash, needs 4GB of VRAM at Q4; the largest, GLM 5.3, needs 41GB. The cheapest device that gets one running is the Radeon RX 7900 XT, with no used-price data yet.

VRAM needed for each GLM version

5 versions
VersionReleasedContext
GLM 4.7 FlashMoE30B (3B active)Jan 19, 2026198K4plus 17 GB of RAMRadeon RX 7900 XT · 20 GB1,935,018
GLM 4.5 AirMoE106B (12B active)Jul 20, 2025128K7plus 56 GB of RAMCMP 170HX 64GB (modded) · 64 GB114,264
GLM 5.3 FlashMoE320B (18B active)Aug 25, 20261M30plus 174 GB of RAMMac Studio M3 Ultra 512GB · 512 GB654,957
GLM 5.2MoE744B (40B active)Jun 16, 20261M41plus 406 GB of RAMMac Studio M3 Ultra 512GB · 512 GB1,134,389
GLM 5.3MoE744B (40B active)Aug 25, 20261M41plus 406 GB of RAMMac Studio M3 Ultra 512GB · 512 GB303,534

Minimum VRAM = Q4_K_M weights + KV cache for an 8K context + 1GB runtime overhead, rounded up. The cheapest device is the first one that fits when sorting by median used price on the current price basis.

About GLM

GLM is published by Z.ai / Zhipu AI under MIT / GLM-5.3 License. This page covers 5 versions ranging from 30B to 744B parameters; structure parameters and quantized sizes are synced weekly from Hugging Face.

VendorZ.ai / Zhipu AI
LicenseMIT / GLM-5.3 License
Official sitehttps://z.ai/
Hugging Face orgzai-org