What Do You Need to Run Gemma Locally? VRAM by Version
Gemma currently has 4 versions you can deploy locally. The smallest, Gemma 4 E4B, needs 7GB of VRAM at Q4; the largest, Gemma 4 31B, needs 26GB. The cheapest device that gets one running is the GeForce RTX 2080 Ti, with no used-price data yet.
VRAM needed for each Gemma version
4 versions| Version | Released | Context | ||||
|---|---|---|---|---|---|---|
| Gemma 4 E4BDense | 8B | Mar 2, 2026 | 128K | 7 | GeForce RTX 2080 Ti— · 11 GB | 4,850,749 |
| Gemma 4 12BDense | 12B | May 23, 2026 | 256K | 11 | GeForce RTX 2080 Ti— · 11 GB | 3,195,490 |
| Gemma 4 26B A4BMoE | 25.8B (4B active) | Mar 11, 2026 | 256K | 6plus 13 GB of RAM | Radeon RX 7900 XT— · 20 GB | 8,211,474 |
| Gemma 4 31BDense | 31.3B | Mar 11, 2026 | 256K | 26 | Arc Pro B65— · 32 GB | 8,332,852 |
Minimum VRAM = Q4_K_M weights + KV cache for an 8K context + 1GB runtime overhead, rounded up. The cheapest device is the first one that fits when sorting by median used price on the current price basis.
About Gemma
Gemma is published by Google DeepMind under Apache-2.0. This page covers 4 versions ranging from 8B to 31.3B parameters; structure parameters and quantized sizes are synced weekly from Hugging Face.
VendorGoogle DeepMind
LicenseApache-2.0
Official sitehttps://ai.google.dev/gemma
Hugging Face orggoogle