This is the most common card question of 2026, and the answer depends on what the card is for: for local LLMs alone, buy a used RTX 3090; if you also game, distrust used hardware, or are limited by case and PSU, buy the RTX 5070 Ti. The two cards have nearly the same bandwidth (936 versus 896 GB/s), so speed differs by under 5%. The difference is memory: 24GB runs 32B dense models at 4-bit and 16GB does not. Prices and trends as of September 5, 2026 are in the tables below.
Key specs side by side
The 3090 wins on 24GB and slightly higher bandwidth; the 5070 Ti wins on power, size, architecture and warranty. Live data:
| Spec | |||
|---|---|---|---|
| VRAM | 24 GB GDDR6X | 16 GB GDDR7 | 16 GB GDDR6X |
| Bandwidth | 936 GB/s | 896 GB/s | 717 GB/s |
| BF16 | 142.3 TFLOPS | 175.8 TFLOPS | 195 TFLOPS |
| INT8 | 285 TOPS | 352 TOPS | 390 TOPS |
| Power | 350 W | 300 W | 320 W |
| Interface | PCIe 4.0 x16 | PCIe 5.0 x16 | PCIe 4.0 x16 |
| Launch date | Sep 24, 2020 | Feb 20, 2025 | Nov 16, 2022 |
| MSRP | $1,499 | $749 | $1,199 |
| eBay (US used) | $1,625 | $1,150 | $1,199 |
| Price per GB | $68 | $72 | $75 |
| Largest dense model (4-bit, 8K) | 32B | 14B | 14B |
| DeepSeek R1 Distill Qwen 14B speed | 68 t/s | 65 t/s | 52 t/s |
What the specs mean for inference:
- Bandwidth is a near tie, which is why speed is a tie. GDDR7's higher clock is cancelled by the 5070 Ti's 256-bit bus against the 3090's 384-bit.
- FP8 / FP4 compute: the 5070 Ti has it, the 3090 does not. It helps vLLM batch serving and FP8 weights; it barely matters for single-stream GGUF in llama.cpp.
- Power: 350W versus 300W on paper, both around 60% of that during inference, a few dollars a year apart.
- Size: the 3090 Founders Edition is three slots and 313 mm, partner cards larger; most 5070 Ti cards are two slots, and an ITX case forces the choice.
What does 24GB run that 16GB cannot?
The extra 8GB is a full tier: 24B to 32B dense models at 4-bit, and 70B models at 2-bit. Catalog versions that only the 3090 runs on one card (4-bit, 8K context):
| Runs on the GeForce RTX 3090 only | VRAM needed GB | Speed on GeForce RTX 3090 t/s |
|---|---|---|
| Gemma 4 26B A4B | 18 | 112 |
| Qwen3.6 27B | 19 | 38 |
| Qwen3.8 27B | 19 | 38 |
| GLM 4.7 Flash | 19 | 144 |
| DeepSeek R1 Distill Qwen 32B | 22 | 32 |
| Qwen3 32B | 22 | 32 |
| Qwen3.6 35B A3B | 22 | 141 |
At 4-bit with a 8K context. 11 models in the catalog run on both cards; 2 run on neither. Cards are compared single-GPU; "GeForce RTX 5070 Ti" gets none of the models above.
This table is the whole case for the 3090. DeepSeek R1 Distill Qwen 32B, Qwen3.8 27B and Qwen3 32B are the models people run most at home in 2026; on the 5070 Ti they drop below 3-bit or spill onto the CPU, losing more than half their speed.
How big is the speed difference?
On models both cards run, under 5%; the 32B models only the 3090 runs land near 32 tokens/s. Estimated decode speed:
| Model | Qwen3 8B | Gemma 4 12B | DeepSeek R1 Distill Qwen 14B | DeepSeek R1 Distill Qwen 32B |
|---|---|---|---|---|
| 114 | 74 | 68 | 32 | |
| 109 | 71 | 65 | Too big | |
| 89 | 58 | 52 | Too big | |
| 117 | 76 | 69 | Too big |
Estimated decode speed at 4-bit with an 8K context, single stream. Values marked * are MoE models running with expert weights in system RAM (70 GB/s assumed).
Public measurements agree: on Llama 3.1 8B Q4_K_M in llama.cpp, community numbers put the 3090 near 120 tokens/s and the 5070 Ti at 110 to 115. The 5070 Ti is about 30% faster at prefill (ingesting a long prompt) thanks to its compute, which you notice with long documents, but the decode speed that sets the feel of a chat is a wash.

The risks of a used 3090
Four risks in order of likelihood: memory chip failure, mining history, burnt power connectors, fan bearings. The first two can be screened out at purchase.
- Memory failure: the 12 GDDR6X chips on the back cool through the backplate and run above 100°C for years before producing artifacts and CUDA errors. Ask the seller to run
memtest_vulkanor the OCCT VRAM test for 30 minutes on video, and run it again yourself on arrival. - Mining: many 3090s mined in 2021 and 2022. Look for washed PCBs, screw marks and worn fans. A mining card is not unbuyable, but it should be cheaper and you should assume the memory has aged hot.
- 12-pin / 8-pin connectors: check for yellowed or blackened plastic.
- Fans: a noisy bearing is a cheap fix.
Where to buy: on eBay choose sellers with 30-day returns and skip the cheap "as-is, no returns" listings.
Price trend: will the 3090 keep falling?
The 3090's used price has been roughly flat through 2025 and 2026, falling less than any new card over the same period. The last 30 days:
| Model | Now | 30-day low | 30-day high | 30-day change | 30d trend |
|---|---|---|---|---|---|
| $1,625 | $1,625 | $1,717 | -5.4% | ||
| $1,150 | $1,150 | $1,223 | -6.0% | ||
| $1,199 | $1,130 | $1,199 | +6.1% |
eBay medians over the last 30 days, one point per 6-hour slot. Change is the current median against the oldest point in the window.
The 3090 is six years old and its price is held up by local-inference demand; the only other 24GB consumer cards are the pricier 4090 and 5090 D V2. There is no reason to wait for a crash. The 5070 Ti is the card worth waiting on: as a current-generation part, its used price a year from now will be well below today's.
Verdict by use case
Models only: 3090. Gaming too, or limited space and power: 5070 Ti. Neither fits: look at the 4080.
- Local LLMs first: RTX 3090. The extra 8GB is a capability tier, not a spec-sheet number.
- Gaming first, models sometimes: RTX 5070 Ti. Much faster at 4K, DLSS 4, lower power, three-year warranty.
- Two-slot case or a 650W PSU: RTX 5070 Ti.
- Tighter budget than the 5070 Ti: a used RTX 4080, the cheapest 16GB CUDA card, see what runs on 16GB.
- Two 3090s: for 70B models above about $3,400, see multi-GPU basics.
