What Do You Need to Run Qwen Locally? VRAM by Version
Qwen currently has 8 versions you can deploy locally. The smallest, Qwen3 8B, needs 7GB of VRAM at Q4; the largest, Qwen3.8 2.4T A95B, needs 30GB. The cheapest device that gets one running is the Tesla V100 16GB, at about $305.
VRAM needed for each Qwen version
eBay
| Version | Released | Context | |||
|---|---|---|---|---|---|
| Qwen3 8BDense | 8.2B | Apr 27, 2025 | 40K | 7 | Tesla V100 16GB$305 · 16 GB |
| Qwen3.5 9BDense | 9.7B | Feb 27, 2026 | 256K | 8 | Tesla V100 16GB$305 · 16 GB |
| Qwen3.6 27BDense | 27.8B | Apr 21, 2026 | 256K | 19 | GeForce RTX 2080 Ti 22GB$551 · 22 GB |
| Qwen3.8 27BDense | 27.8B | Aug 5, 2026 | 256K | 19 | GeForce RTX 2080 Ti 22GB$551 · 22 GB |
| Qwen3 32BDense | 32.8B | Apr 27, 2025 | 40K | 22 | GeForce RTX 2080 Ti 22GB$551 · 22 GB |
| Qwen3.6 35B A3BMoE | 36B (3B active) | Apr 15, 2026 | 256K | 4plus 19 GB of RAM | Tesla V100 16GB$305 · 16 GB · Experts offloaded |
| Qwen3.8 Flash NextMoE | 176B (6B active) | Aug 24, 2026 | 256K | 4plus 97 GB of RAM | Tesla V100 16GB$305 · 16 GB · Experts offloaded |
| Qwen3.8 2.4T A95BMoE | 2,446B (95B active) | Aug 8, 2026 | 256K | 30plus 1,337 GB of RAM | Tesla V100 32GB$720 · 32 GB · Experts offloaded |
Minimum VRAM = Q4_K_M weights + KV cache for an 8K context + 1GB runtime overhead, rounded up. The cheapest device is the first one that fits when sorting by median price on the current price basis.
About Qwen
Qwen is published by Alibaba under Apache-2.0 / Qwen Community License 1.0 / Qwen3.8-Max License. This page covers 8 versions ranging from 8.2B to 2,446B parameters. New official releases are scanned weekly; architecture parameters and sourced or estimated weight sizes are reviewed before publication.
VendorAlibaba
LicenseApache-2.0 / Qwen Community License 1.0 / Qwen3.8-Max License
Official sitehttps://qwen.ai/
Hugging Face orgQwen