GPU price index
Lowest price per GB
Lowest price per 100 GB/s
Lowest price per TOPS
Latest release
Biggest 24h drops
No price moves in the last 24h
Biggest 24h rises
No price moves in the last 24h
Most listings
- 1GeForce RTX 5070 Ti$1,57610 listings
- 2GeForce RTX 5070$95610 listings
- 3Radeon RX 7800 XT$96610 listings
- 4GeForce RTX 3090$2,5209 listings
- 5GeForce RTX 4090$4,9059 listings
Latest releases: minimum VRAM
- 1
Hy4 preview770B18 GB - 2
GLM 5.3744B41 GB - 3
GLM 5.3 Flash320B30 GB - 4
Qwen3.8 Flash Next176B4 GB - 5
DeepSeek V4 Pro1,600B17 GB
Local LLM GPU Ranking
AmazonJP
| # | Model | 30d trend | Ecosystem | Buy | ||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | $7,4899 listings | +0.5%6h | 32 | 1,792 | 419 | 838 | 575 | $234 | $418 | $8.9 | CUDA | Check prices | ||
| 2 | $4,9059 listings | +0.3%6h | 24 | 1,008 | 330.3 | 660.6 | 450 | $204 | $487 | $7.4 | CUDA | Check prices | ||
| 3 | $1,8684 listings | 0.0%6h | 16 | 960 | 225.1 | 450.2 | 360 | $117 | $195 | $4.15 | CUDA | Check prices | ||
| 4 | $2,1878 listings | -2.3%6h | 16 | 717 | 195 | 389.9 | 320 | $137 | $305 | $5.6 | CUDA | Check prices | ||
| 5 | $884Using previous price8 listings | -11.3%6h | 16 | 640 | 195 | 389 | 304 | $55 | $138 | $2.27 | ROCm | Check prices | ||
| 6 | $2,2446 listings | 0.0%6h | 32 | 640 | 191 | 383 | 300 | $70 | $351 | $5.9 | ROCm | Check prices | ||
| 7 | $2,0743 listings | 0.0%6h | 32 | 608 | 183.5 | 367 | 230 | $65 | $341 | $5.7 | oneAPI | Check prices | ||
| 8 | $1,576Using previous price10 listings | 0.0%6h | 16 | 896 | 175.8 | 351.5 | 300 | $99 | $176 | $4.48 | CUDA | Check prices | ||
| 9 | $3,9565 listings | -1.0%6h | 24 | 672 | 161.3 | 322.5 | 145 | $165 | $589 | $12.3 | CUDA | Check prices | ||
| 10 | $2,520Using previous price9 listings | -2.5%6h | 24 | 936 | 142.3 | 284.7 | 350 | $105 | $269 | $8.9 | CUDA | Check prices | ||
| 11 | $8,2233 listings | -0.7%6h | 128 | 273 | 125 | 250 | 240 | $64 | $3,012 | $32.9 | CUDA | Check prices | ||
| 12 | $956Using previous price10 listings | 0.0%6h | 12 | 672 | 123.5 | 247 | 250 | $80 | $142 | $3.87 | CUDA | Check prices | ||
| 13 | $1,1643 listings | 0.0%6h | 24 | 456 | 98.5 | 197 | 200 | $48.5 | $255 | $5.9 | oneAPI | Check prices | ||
| 14 | $924Using previous price9 listings | 0.0%6h | 16 | 448 | 94.9 | 189.8 | 180 | $58 | $206 | $4.87 | CUDA | Check prices | ||
| 15 | $1,7389 listings | 0.0%6h | 24 | 960 | 123 | 123 | 355 | $72 | $181 | $14.1 | ROCm | Check prices | ||
| 16 | $1,2449 listings | +6.2%6h | 20 | 800 | 103 | 103 | 315 | $62 | $155 | $12.1 | ROCm | Check prices | ||
| 17 | $1,4044 listings | 0.0%6h | 16 | 576 | 47.3 | 94.6 | 335 | $88 | $244 | $14.8 | ROCm | Check prices | ||
| 18 | $96610 listings | -1.8%6h | 16 | 624 | 74.7 | 74.7 | 263 | $60 | $155 | $12.9 | ROCm | Check prices | ||
| 19 | $3,9136 listings | 0.0%6h | 128 | 256 | 59.4 | 59.4 | 120 | $30.6 | $1,528 | $66 | ROCm | Check prices | ||
| 20 | $6594 listings | -0.6%6h | 16 | 900 | 112 | — | 250 | $41.2 | $73 | — | CUDA | Check prices | ||
| 21 | $1,1413 listings | 0.0%6h | 32 | 900 | 112 | — | 250 | $35.7 | $127 | — | CUDA | Check prices |
Choose hardware for local LLM deployment
Find the right GPU for local LLMs. Compare VRAM, GPU prices and model memory requirements to choose hardware that fits your models and budget.
What is VRAM, and how much do you need?
VRAM is the fast memory on a graphics card. Local LLM inference needs space for model weights, the context (KV) cache and runtime buffers. Quantization can reduce weight memory, while longer prompts and more simultaneous users need additional memory. A model fitting in VRAM does not guarantee a particular speed.
Compare model VRAM requirementsWhat runs on 16GB of VRAM?Plan a local deployment
Choose a model and quantization, set the context length, and check both GPU memory and system RAM. CPU or expert offloading can make a large model fit with less VRAM, but it depends on RAM bandwidth and can be much slower. Check CUDA, ROCm or Metal support before buying hardware.
Local LLM hardware and deployment guidesRead the rankings and price data
Compare VRAM capacity first, then bandwidth, power and software support. Our prices are median asking prices, not completed sales; sample counts, observation dates and stale-price labels help you judge their usefulness. INT8 throughput alone is not an LLM benchmark.
Price sources and calculation methodsPrices are median asking prices, not sold prices · Methodology