GPUs tracked39Price history3 dayssince Sep 3, 2026Price points (24h)11,753 pointsBiggest 24h dropBiggest 24h riseData updated Sep 5, 2026

What Do You Need to Run Qwen Locally? VRAM by Version

Qwen currently has 6 versions you can deploy locally. The smallest, Qwen3 8B, needs 7GB of VRAM at Q4; the largest, Qwen3.6 35B A3B, needs 4GB. The cheapest device that gets one running is the GeForce RTX 2080 Ti, with no used-price data yet.

VRAM needed for each Qwen version

6 versions
VersionReleasedContext
Qwen3 8BDense8.2BApr 27, 202540K7GeForce RTX 2080 Ti · 11 GB13,232,997
Qwen3.5 9BDense9.7BFeb 27, 2026256K8GeForce RTX 2080 Ti · 11 GB12,375,273
Qwen3.6 27BDense27.8BApr 21, 2026256K19Radeon RX 7900 XT · 20 GB5,366,737
Qwen3.8 27BDense27.8BAug 5, 2026256K19Radeon RX 7900 XT · 20 GB5,739,341
Qwen3 32BDense32.8BApr 27, 202540K22Arc Pro B60 · 24 GB5,024,271
Qwen3.6 35B A3BMoE36B (3B active)Apr 15, 2026256K4plus 19 GB of RAMArc Pro B60 · 24 GB4,546,612

Minimum VRAM = Q4_K_M weights + KV cache for an 8K context + 1GB runtime overhead, rounded up. The cheapest device is the first one that fits when sorting by median used price on the current price basis.

About Qwen

Qwen is published by Alibaba under Apache-2.0. This page covers 6 versions ranging from 8.2B to 36B parameters; structure parameters and quantized sizes are synced weekly from Hugging Face.

VendorAlibaba
LicenseApache-2.0
Official sitehttps://qwen.ai/
Hugging Face orgQwen