GPUs tracked45Price history8 dayssince Sep 3, 2026Price points (24h)6,955 pointsBiggest 24h dropBiggest 24h riseData updated

What Do You Need to Run Qwen Locally? VRAM by Version

Qwen currently has 8 versions you can deploy locally. The smallest, Qwen3 8B, needs 7GB of VRAM at Q4; the largest, Qwen3.8 2.4T A95B, needs 30GB. The cheapest device that gets one running is the Tesla V100 16GB, at about $659.

VRAM needed for each Qwen version

AmazonJP
8 versions
VersionReleasedContext
Qwen3 8BDense8.2BApr 27, 202540K7Tesla V100 16GB$659 · 16 GB
Qwen3.5 9BDense9.7BFeb 27, 2026256K8Tesla V100 16GB$659 · 16 GB
Qwen3.6 27BDense27.8BApr 21, 2026256K19Tesla V100 32GB$1,141 · 32 GB
Qwen3.8 27BDense27.8BAug 5, 2026256K19Tesla V100 32GB$1,141 · 32 GB
Qwen3 32BDense32.8BApr 27, 202540K22Tesla V100 32GB$1,141 · 32 GB
Qwen3.6 35B A3BMoE36B (3B active)Apr 15, 2026256K4plus 19 GB of RAMTesla V100 16GB$659 · 16 GB · Experts offloaded
Qwen3.8 Flash NextMoE176B (6B active)Aug 24, 2026256K4plus 97 GB of RAMTesla V100 16GB$659 · 16 GB · Experts offloaded
Qwen3.8 2.4T A95BMoE2,446B (95B active)Aug 8, 2026256K30plus 1,337 GB of RAMTesla V100 32GB$1,141 · 32 GB · Experts offloaded

Minimum VRAM = Q4_K_M weights + KV cache for an 8K context + 1GB runtime overhead, rounded up. The cheapest device is the first one that fits when sorting by median price on the current price basis.

About Qwen

Qwen is published by Alibaba under Apache-2.0 / Qwen Community License 1.0 / Qwen3.8-Max License. This page covers 8 versions ranging from 8.2B to 2,446B parameters. New official releases are scanned weekly; architecture parameters and sourced or estimated weight sizes are reviewed before publication.

VendorAlibaba
LicenseApache-2.0 / Qwen Community License 1.0 / Qwen3.8-Max License
Official sitehttps://qwen.ai/
Hugging Face orgQwen