This is the most common card question of 2026, and the answer depends on what the card is for: for local LLMs alone, buy a used RTX 3090; if you also game, distrust used hardware, or are limited by case and PSU, buy the RTX 5070 Ti. The two cards have nearly the same bandwidth (936 versus 896 GB/s), so speed differs by under 5%. The difference is memory: 24GB runs 32B dense models at 4-bit and 16GB does not. Prices and trends as of September 5, 2026 are in the tables below.

Key specs side by side

The 3090 wins on 24GB and slightly higher bandwidth; the 5070 Ti wins on power, size, architecture and warranty. Live data:

SpecGeForce RTX 3090GeForce RTX 5070 TiGeForce RTX 4080
VRAM24 GB GDDR6X16 GB GDDR716 GB GDDR6X
Bandwidth936 GB/s896 GB/s717 GB/s
BF16142.3 TFLOPS175.8 TFLOPS195 TFLOPS
INT8285 TOPS352 TOPS390 TOPS
Power350 W300 W320 W
InterfacePCIe 4.0 x16PCIe 5.0 x16PCIe 4.0 x16
Launch dateSep 24, 2020Feb 20, 2025Nov 16, 2022
MSRP$1,499$749$1,199
eBay (US used)$1,625$1,150$1,199
Price per GB$68$72$75
Largest dense model (4-bit, 8K)32B14B14B
DeepSeek R1 Distill Qwen 14B speed68 t/s65 t/s52 t/s

What the specs mean for inference:

  • Bandwidth is a near tie, which is why speed is a tie. GDDR7's higher clock is cancelled by the 5070 Ti's 256-bit bus against the 3090's 384-bit.
  • FP8 / FP4 compute: the 5070 Ti has it, the 3090 does not. It helps vLLM batch serving and FP8 weights; it barely matters for single-stream GGUF in llama.cpp.
  • Power: 350W versus 300W on paper, both around 60% of that during inference, a few dollars a year apart.
  • Size: the 3090 Founders Edition is three slots and 313 mm, partner cards larger; most 5070 Ti cards are two slots, and an ITX case forces the choice.

What does 24GB run that 16GB cannot?

The extra 8GB is a full tier: 24B to 32B dense models at 4-bit, and 70B models at 2-bit. Catalog versions that only the 3090 runs on one card (4-bit, 8K context):

Runs on the GeForce RTX 3090 onlyVRAM needed GBSpeed on GeForce RTX 3090 t/s
Gemma 4 26B A4B18112
Qwen3.6 27B1938
Qwen3.8 27B1938
GLM 4.7 Flash19144
DeepSeek R1 Distill Qwen 32B2232
Qwen3 32B2232
Qwen3.6 35B A3B22141

At 4-bit with a 8K context. 11 models in the catalog run on both cards; 2 run on neither. Cards are compared single-GPU; "GeForce RTX 5070 Ti" gets none of the models above.

This table is the whole case for the 3090. DeepSeek R1 Distill Qwen 32B, Qwen3.8 27B and Qwen3 32B are the models people run most at home in 2026; on the 5070 Ti they drop below 3-bit or spill onto the CPU, losing more than half their speed.

How big is the speed difference?

On models both cards run, under 5%; the 32B models only the 3090 runs land near 32 tokens/s. Estimated decode speed:

ModelQwen3 8BGemma 4 12BDeepSeek R1 Distill Qwen 14BDeepSeek R1 Distill Qwen 32B
GeForce RTX 3090114746832
GeForce RTX 5070 Ti1097165Too big
GeForce RTX 4080895852Too big
GeForce RTX 50801177669Too big

Estimated decode speed at 4-bit with an 8K context, single stream. Values marked * are MoE models running with expert weights in system RAM (70 GB/s assumed).

Public measurements agree: on Llama 3.1 8B Q4_K_M in llama.cpp, community numbers put the 3090 near 120 tokens/s and the 5070 Ti at 110 to 115. The 5070 Ti is about 30% faster at prefill (ingesting a long prompt) thanks to its compute, which you notice with long documents, but the decode speed that sets the feel of a chat is a wash.

Illustration: hands holding a used graphics card under a desk lamp to inspect the backplate and power connector
Illustration: hands holding a used graphics card under a desk lamp to inspect the backplate and power connector

The risks of a used 3090

Four risks in order of likelihood: memory chip failure, mining history, burnt power connectors, fan bearings. The first two can be screened out at purchase.

  1. Memory failure: the 12 GDDR6X chips on the back cool through the backplate and run above 100°C for years before producing artifacts and CUDA errors. Ask the seller to run memtest_vulkan or the OCCT VRAM test for 30 minutes on video, and run it again yourself on arrival.
  2. Mining: many 3090s mined in 2021 and 2022. Look for washed PCBs, screw marks and worn fans. A mining card is not unbuyable, but it should be cheaper and you should assume the memory has aged hot.
  3. 12-pin / 8-pin connectors: check for yellowed or blackened plastic.
  4. Fans: a noisy bearing is a cheap fix.

Where to buy: on eBay choose sellers with 30-day returns and skip the cheap "as-is, no returns" listings.

Price trend: will the 3090 keep falling?

The 3090's used price has been roughly flat through 2025 and 2026, falling less than any new card over the same period. The last 30 days:

ModelNow30-day low30-day high30-day change30d trend
GeForce RTX 3090$1,625$1,625$1,717-5.4%
GeForce RTX 5070 Ti$1,150$1,150$1,223-6.0%
GeForce RTX 4080$1,199$1,130$1,199+6.1%

eBay medians over the last 30 days, one point per 6-hour slot. Change is the current median against the oldest point in the window.

The 3090 is six years old and its price is held up by local-inference demand; the only other 24GB consumer cards are the pricier 4090 and 5090 D V2. There is no reason to wait for a crash. The 5070 Ti is the card worth waiting on: as a current-generation part, its used price a year from now will be well below today's.

Verdict by use case

Models only: 3090. Gaming too, or limited space and power: 5070 Ti. Neither fits: look at the 4080.

  • Local LLMs first: RTX 3090. The extra 8GB is a capability tier, not a spec-sheet number.
  • Gaming first, models sometimes: RTX 5070 Ti. Much faster at 4K, DLSS 4, lower power, three-year warranty.
  • Two-slot case or a 650W PSU: RTX 5070 Ti.
  • Tighter budget than the 5070 Ti: a used RTX 4080, the cheapest 16GB CUDA card, see what runs on 16GB.
  • Two 3090s: for 70B models above about $3,400, see multi-GPU basics.