GPUs tracked45Price history8 dayssince Sep 3, 2026Price points (24h)6,955 pointsBiggest 24h dropMac Studio M4 Max 128GB-6.7%Biggest 24h riseGeForce RTX 5070 Ti+21.0%Data updated

What Do You Need to Run GLM Locally? VRAM by Version

GLM currently has 5 versions you can deploy locally. The smallest, GLM 4.7 Flash, needs 4GB of VRAM at Q4; the largest, GLM 5.3, needs 41GB. The cheapest device that gets one running is the Tesla V100 16GB, at about $131.

VRAM needed for each GLM version

Xianyu
5 versions
VersionReleasedContext
GLM 4.7 FlashMoE30B (3B active)Jan 19, 2026198K4plus 17 GB of RAMTesla V100 16GB$131 · 16 GB · Experts offloaded
GLM 4.5 AirMoE106B (12B active)Jul 20, 2025128K7plus 56 GB of RAMTesla V100 16GB$131 · 16 GB · Experts offloaded
GLM 5.3 FlashMoE320B (18B active)Aug 25, 20261M30plus 174 GB of RAMInstinct MI50 32GB$405 · 32 GB · Experts offloaded
GLM 5.2MoE744B (40B active)Jun 16, 20261M41plus 406 GB of RAMCMP 170HX 64GB$2,089 · 64 GB · Experts offloaded
GLM 5.3MoE744B (40B active)Aug 25, 20261M41plus 406 GB of RAMCMP 170HX 64GB$2,089 · 64 GB · Experts offloaded

Minimum VRAM = Q4_K_M weights + KV cache for an 8K context + 1GB runtime overhead, rounded up. The cheapest device is the first one that fits when sorting by median price on the current price basis.

About GLM

GLM is published by Z.ai / Zhipu AI under MIT / GLM-5.3 License. This page covers 5 versions ranging from 30B to 744B parameters. New official releases are scanned weekly; architecture parameters and sourced or estimated weight sizes are reviewed before publication.

VendorZ.ai / Zhipu AI
LicenseMIT / GLM-5.3 License
Official sitehttps://z.ai/
Hugging Face orgzai-org