What Do You Need to Run Qwen Locally? VRAM by Version
Qwen currently has 6 versions you can deploy locally. The smallest, Qwen3 8B, needs 7GB of VRAM at Q4; the largest, Qwen3.6 35B A3B, needs 4GB. The cheapest device that gets one running is the GeForce RTX 2080 Ti, with no used-price data yet.
VRAM needed for each Qwen version
6 versions| Version | Released | Context | ||||
|---|---|---|---|---|---|---|
| Qwen3 8BDense | 8.2B | Apr 27, 2025 | 40K | 7 | GeForce RTX 2080 Ti— · 11 GB | 13,232,997 |
| Qwen3.5 9BDense | 9.7B | Feb 27, 2026 | 256K | 8 | GeForce RTX 2080 Ti— · 11 GB | 12,375,273 |
| Qwen3.6 27BDense | 27.8B | Apr 21, 2026 | 256K | 19 | Radeon RX 7900 XT— · 20 GB | 5,366,737 |
| Qwen3.8 27BDense | 27.8B | Aug 5, 2026 | 256K | 19 | Radeon RX 7900 XT— · 20 GB | 5,739,341 |
| Qwen3 32BDense | 32.8B | Apr 27, 2025 | 40K | 22 | Arc Pro B60— · 24 GB | 5,024,271 |
| Qwen3.6 35B A3BMoE | 36B (3B active) | Apr 15, 2026 | 256K | 4plus 19 GB of RAM | Arc Pro B60— · 24 GB | 4,546,612 |
Minimum VRAM = Q4_K_M weights + KV cache for an 8K context + 1GB runtime overhead, rounded up. The cheapest device is the first one that fits when sorting by median used price on the current price basis.
About Qwen
Qwen is published by Alibaba under Apache-2.0. This page covers 6 versions ranging from 8.2B to 36B parameters; structure parameters and quantized sizes are synced weekly from Hugging Face.
VendorAlibaba
LicenseApache-2.0
Official sitehttps://qwen.ai/
Hugging Face orgQwen