GPUs tracked45Price history8 dayssince Sep 3, 2026Price points (24h)6,955 pointsBiggest 24h dropMac Studio M4 Max 128GB-6.7%Biggest 24h riseGeForce RTX 5070 Ti+21.0%Data updated

What Do You Need to Run Qwen Locally? VRAM by Version

Qwen currently has 8 versions you can deploy locally. The smallest, Qwen3 8B, needs 7GB of VRAM at Q4; the largest, Qwen3.8 2.4T A95B, needs 30GB. The cheapest device that gets one running is the Tesla V100 16GB, at about $131.

VRAM needed for each Qwen version

Xianyu
8 versions
VersionReleasedContext
Qwen3 8BDense8.2BApr 27, 202540K7Tesla V100 16GB$131 · 16 GB
Qwen3.5 9BDense9.7BFeb 27, 2026256K8Tesla V100 16GB$131 · 16 GB
Qwen3.6 27BDense27.8BApr 21, 2026256K19GeForce RTX 2080 Ti 22GB$387 · 22 GB
Qwen3.8 27BDense27.8BAug 5, 2026256K19GeForce RTX 2080 Ti 22GB$387 · 22 GB
Qwen3 32BDense32.8BApr 27, 202540K22GeForce RTX 2080 Ti 22GB$387 · 22 GB
Qwen3.6 35B A3BMoE36B (3B active)Apr 15, 2026256K4plus 19 GB of RAMTesla V100 16GB$131 · 16 GB · Experts offloaded
Qwen3.8 Flash NextMoE176B (6B active)Aug 24, 2026256K4plus 97 GB of RAMTesla V100 16GB$131 · 16 GB · Experts offloaded
Qwen3.8 2.4T A95BMoE2,446B (95B active)Aug 8, 2026256K30plus 1,337 GB of RAMInstinct MI50 32GB$405 · 32 GB · Experts offloaded

Minimum VRAM = Q4_K_M weights + KV cache for an 8K context + 1GB runtime overhead, rounded up. The cheapest device is the first one that fits when sorting by median price on the current price basis.

About Qwen

Qwen is published by Alibaba under Apache-2.0 / Qwen Community License 1.0 / Qwen3.8-Max License. This page covers 8 versions ranging from 8.2B to 2,446B parameters. New official releases are scanned weekly; architecture parameters and sourced or estimated weight sizes are reviewed before publication.

VendorAlibaba
LicenseApache-2.0 / Qwen Community License 1.0 / Qwen3.8-Max License
Official sitehttps://qwen.ai/
Hugging Face orgQwen