What Do You Need to Run Llama Locally? VRAM by Version
Llama currently has 4 versions you can deploy locally. The smallest, Llama 3.2 3B, needs 4GB of VRAM at Q4; the largest, Llama 4 Scout 17B 16E, needs 10GB. The cheapest device that gets one running is the GeForce RTX 2080 Ti, with no used-price data yet.
VRAM needed for each Llama version
4 versions| Version | Released | Context | ||||
|---|---|---|---|---|---|---|
| Llama 3.2 3BDense | 3.2B | Sep 18, 2024 | 128K | 4 | GeForce RTX 2080 Ti— · 11 GB | 1,419,885 |
| Llama 3.1 8BDense | 8B | Jul 18, 2024 | 128K | 7 | GeForce RTX 2080 Ti— · 11 GB | 5,734,979 |
| Llama 3.3 70BDense | 70.6B | Nov 26, 2024 | 128K | 43 | GeForce RTX 4090 48GB (modded)— · 48 GB | 815,592 |
| Llama 4 Scout 17B 16EMoE | 108.6B (17B active) | Apr 2, 2025 | 10M | 10plus 55 GB of RAM | CMP 170HX 64GB (modded)— · 64 GB | 174,923 |
Minimum VRAM = Q4_K_M weights + KV cache for an 8K context + 1GB runtime overhead, rounded up. The cheapest device is the first one that fits when sorting by median used price on the current price basis.
About Llama
Llama is published by Meta AI under Llama 3.2 Community License / Llama 3.1 Community License / Llama 3.3 Community License / Llama 4 Community License. This page covers 4 versions ranging from 3.2B to 108.6B parameters; structure parameters and quantized sizes are synced weekly from Hugging Face.