GeForce RTX 5070 Price & Used Prices: Which Local LLMs Can It Run?
The GeForce RTX 5070 has 12GB of GDDR7 memory with 672 GB/s of bandwidth. No used-price data yet. At Q4 quantization it runs models up to 14B parameters at roughly 52 tokens/s.
GeForce RTX 5070 used price history
Full GeForce RTX 5070 specifications
What LLMs can the GeForce RTX 5070 run?
Versions released since 2026 with up to 50B parameters| Model | 2-bit | 4-bit | 8-bit | ||||||
|---|---|---|---|---|---|---|---|---|---|
| Fits | Speed t/s | Context | Fits | Speed t/s | Context | Fits | Speed t/s | Context | |
| GemmaGemma 4 E4B | ✓ | 125.6 | 32K | ✓ | 89.4 | 32K | ✓ | 51.2 | 32K |
| QwenQwen3.5 9B | ✓ | 103.5 | 32K | ✓ | 73.8 | 32K | ✗ | — | — |
| GemmaGemma 4 12B | ✓ | 73.3 | 8K | ✓ | 54.1 | 8K | ✗ | — | — |
| GemmaGemma 4 26B A4B | Offload | 51.7 | — | Offload | 38.5 | — | Offload | 23.2 | — |
| QwenQwen3.6 27B | ✗ | — | — | ✗ | — | — | ✗ | — | — |
| QwenQwen3.8 27B | ✗ | — | — | ✗ | — | — | ✗ | — | — |
| GLMGLM 4.7 Flash | Offload | 51.6 | — | Offload | 37.3 | — | Offload | 21.7 | — |
| GemmaGemma 4 31B | ✗ | — | — | ✗ | — | — | ✗ | — | — |
| QwenQwen3.6 35B A3B | Offload | 69.1 | — | Offload | 51.6 | — | Offload | 31.2 | — |
Fits when weights + KV cache + runtime overhead is at or below usable memory, single card. "Offload" on MoE models means it runs with expert weights in system RAM (only attention layers and KV cache stay in VRAM); that speed assumes 70 GB/s RAM bandwidth. The context column is the largest tier that fits.
Prices by platform
Buy links may carry affiliate tags; they do not change the price you pay.
FAQ
How much VRAM does the GeForce RTX 5070 have?
The GeForce RTX 5070 has 12GB of GDDR7 memory with 672 GB/s of bandwidth.
Can the GeForce RTX 5070 run a 70B model?
Not on a single card. At Q4_K_M with an 8K context a 70B dense model needs about 43GB, more than the GeForce RTX 5070's 12GB.
How much does a used GeForce RTX 5070 cost?
No used-price data yet; it will appear here once the feed is live.