GPUs tracked45Price history8 dayssince Sep 3, 2026Price points (24h)7,968 pointsBiggest 24h dropGeForce RTX 5070 Ti-6.3%Biggest 24h riseRTX PRO 4000 Blackwell+7.7%Data updated

What Do You Need to Run Llama Locally? VRAM by Version

Llama currently has 4 versions you can deploy locally. The smallest, Llama 3.2 3B, needs 4GB of VRAM at Q4; the largest, Llama 4 Scout 17B 16E, needs 10GB. The cheapest device that gets one running is the Tesla V100 16GB, at about $305.

VRAM needed for each Llama version

eBay
4 versions
VersionReleasedContext
Llama 3.2 3BDense3.2BSep 18, 2024128K4Tesla V100 16GB$305 · 16 GB
Llama 3.1 8BDense8BJul 18, 2024128K7Tesla V100 16GB$305 · 16 GB
Llama 3.3 70BDense70.6BNov 26, 2024128K43Radeon PRO W7900$3,495 · 48 GB
Llama 4 Scout 17B 16EMoE108.6B (17B active)Apr 2, 202510M10plus 55 GB of RAMTesla V100 16GB$305 · 16 GB · Experts offloaded

Minimum VRAM = Q4_K_M weights + KV cache for an 8K context + 1GB runtime overhead, rounded up. The cheapest device is the first one that fits when sorting by median price on the current price basis.

About Llama

Llama is published by Meta AI under Llama 3.2 Community License / Llama 3.1 Community License / Llama 3.3 Community License / Llama 4 Community License. This page covers 4 versions ranging from 3.2B to 108.6B parameters. New official releases are scanned weekly; architecture parameters and sourced or estimated weight sizes are reviewed before publication.

VendorMeta AI
LicenseLlama 3.2 Community License / Llama 3.1 Community License / Llama 3.3 Community License / Llama 4 Community License
Hugging Face orgmeta-llama