What Do You Need to Run Mistral Locally? VRAM by Version
Mistral currently has 5 versions you can deploy locally. The smallest, Mistral 7B v0.3, needs 7GB of VRAM at Q4; the largest, Mistral Small 3.2 24B, needs 16GB. The cheapest device that gets one running is the GeForce RTX 2080 Ti, with no used-price data yet.
VRAM needed for each Mistral version
5 versions| Version | Released | Context | ||||
|---|---|---|---|---|---|---|
| Mistral 7B v0.3Dense | 7.3B | May 22, 2024 | 32K | 7 | GeForce RTX 2080 Ti— · 11 GB | 2,648,636 |
| Ministral 3 8BDense | 8.9B | Oct 31, 2025 | 256K | 8 | GeForce RTX 2080 Ti— · 11 GB | 142,610 |
| Mistral NemoDense | 12.3B | Jul 17, 2024 | 128K | 10 | GeForce RTX 2080 Ti— · 11 GB | 358,821 |
| Ministral 3 14BDense | 14B | Oct 31, 2025 | 256K | 11 | GeForce RTX 2080 Ti— · 11 GB | 266,145 |
| Mistral Small 3.2 24BDense | 24B | Jun 19, 2025 | 128K | 16 | GeForce RTX 4080— · 16 GB | 126,611 |
Minimum VRAM = Q4_K_M weights + KV cache for an 8K context + 1GB runtime overhead, rounded up. The cheapest device is the first one that fits when sorting by median used price on the current price basis.
About Mistral
Mistral is published by Mistral AI under Apache-2.0. This page covers 5 versions ranging from 7.3B to 24B parameters; structure parameters and quantized sizes are synced weekly from Hugging Face.
VendorMistral AI
LicenseApache-2.0
Official sitehttps://mistral.ai/
Hugging Face orgmistralai