What Do You Need to Run Gemma Locally? VRAM by Version
Gemma currently has 4 versions you can deploy locally. The smallest, Gemma 4 E4B, needs 7GB of VRAM at Q4; the largest, Gemma 4 31B, needs 26GB. The cheapest device that gets one running is the Tesla V100 16GB, at about $659.
VRAM needed for each Gemma version
AmazonJP
| Version | Released | Context | |||
|---|---|---|---|---|---|
| Gemma 4 E4BDense | 8B | Mar 2, 2026 | 128K | 7 | Tesla V100 16GB$659 · 16 GB |
| Gemma 4 12BDense | 12B | May 23, 2026 | 256K | 11 | Tesla V100 16GB$659 · 16 GB |
| Gemma 4 26B A4BMoE | 25.8B (4B active) | Mar 11, 2026 | 256K | 6plus 13 GB of RAM | Tesla V100 16GB$659 · 16 GB · Experts offloaded |
| Gemma 4 31BDense | 31.3B | Mar 11, 2026 | 256K | 26 | Tesla V100 32GB$1,141 · 32 GB |
Minimum VRAM = Q4_K_M weights + KV cache for an 8K context + 1GB runtime overhead, rounded up. The cheapest device is the first one that fits when sorting by median price on the current price basis.
About Gemma
Gemma is published by Google DeepMind under Apache-2.0. This page covers 4 versions ranging from 8B to 31.3B parameters. New official releases are scanned weekly; architecture parameters and sourced or estimated weight sizes are reviewed before publication.
VendorGoogle DeepMind
LicenseApache-2.0
Official sitehttps://ai.google.dev/gemma
Hugging Face orggoogle