GPUs tracked56Price history27 daysPrice points (24h)34,228 pointsBiggest 24h dropGeForce RTX 4060 Ti 16GB-16.4%Biggest 24h riseRadeon RX 7800 XT+19.1%Updated
$ USD
English

GeForce RTX 4090D 48GB Price & Used Prices: Which Local LLMs Can It Run?

The GeForce RTX 4090D 48GB has 48GB of GDDR6X memory with 1,008 GB/s of bandwidth. No price data yet. At 4-bit quantization and 8K context it fits a dense model of up to about 79B parameters, at roughly 15 tokens/s.

GeForce RTX 4090D 48GB price history

Price updated —

Not enough data to show a trend

The median price trend will appear once enough valid quotes are available.

Buy links may carry affiliate tags; they do not change the price you pay.

Full GeForce RTX 4090D 48GB specifications

  • Memory48 GB · GDDR6X
  • Bandwidth1,008 GB/s
  • FP16 / BF16294.3 TFLOPS
  • FP8 / INT8588.5 / 588.5 TOPS
Bus width
384-bit
Architecture
Ada Lovelace (AD102)
Cores
14,592
Tensor cores
456
Ecosystem
CUDA
Interface
PCIe 4.0 x16
NVLink
Not supported
Power
425 W
Power connector
16-pin 12VHPWR
Slots
2
Type
Consumer
Launch date
Dec 28, 2023

What LLMs can the GeForce RTX 4090D 48GB run?

Models up to 50B parameters on a single GeForce RTX 4090D 48GB, versions released since 2026 only

ModelSizeRuns?Speed t/sMax context
Gemma 4 E4B8B5 GBRuns130.7128K
Qwen3.5 9B9.7B5.7 GBRuns115.3256K
Gemma 4 12B12B7.1 GBRuns93.5256K
Gemma 4 26B A4B25.2B (3.8B active)16.9 GBRuns130.1256K
Qwen3.6 27B27.8B16.8 GBRuns42256K
Qwen3.8 27B27.8B16.5 GBRuns42.7256K
GLM 4.7 Flash30B (3B active)18.3 GBRuns147.9198K
Gemma 4 31B30.7B18.3 GBRuns37.7256K
Qwen3.6 35B A3B36B (3B active)22.1 GBRuns151.5256K

RunsFits fully in VRAM; speed is estimated from VRAM bandwidth.

Needs RAM offloadMoE only: expert weights sit in system RAM, so speed is estimated from 70 GB/s RAM bandwidth.

Too bigWon’t fit on one card at this quantization, even with RAM offload.

Dual, 4x and 8x GeForce RTX 4090D 48GB for LLMs: which models fit?

Cards needed for 50B+ models released since 2026, with vLLM tensor parallelism at 32K context

ModelSizeCardsTotal priceTotal VRAMSpeed t/sTotal power
Qwen3.8 Flash Next176B (6B active)111.3 GB4 cardsNo price192 GB103.61,700 W
DeepSeek V4 Flash284B (13B active)155.1 GB4 cardsNo price192 GB731,700 W
Hy3295B (21B active)182.2 GB8 cardsNo price384 GB353,400 W
GLM 5.3 Flash320B (18B active)199.7 GB8 cardsNo price384 GB51.53,400 W

Card counts are 1 / 2 / 4 / 8, each using 90% of its VRAM. Speed is based on one card’s bandwidth; price covers GPUs only. No NVLink; cards talk over PCIe. Whether it runs also depends on vLLM support.

Other markets

eBay · United States
N/A
No listings
View
Xianyu · Mainland China
$3,526
9 listings
View

Buy links may carry affiliate tags; they do not change the price you pay.

FAQ

How large a model can the GeForce RTX 4090D 48GB run?

A single GeForce RTX 4090D 48GB has 48GB of GDDR6X VRAM. At 4-bit quantization and 8K context it fits a dense model of up to about 79B parameters, or about 43B at 8-bit.

Can the GeForce RTX 4090D 48GB run a 70B model?

Yes. At 4-bit with an 8K context a 70B dense model needs about 43GB, which fits in the GeForce RTX 4090D 48GB's 48GB.

What are the best LLMs to run on the GeForce RTX 4090D 48GB?

Among models released in 2026, the largest dense model that fits entirely in VRAM at 4-bit is Gemma 4 31B, at an estimated 37.7 tokens/s. Among MoE models, Qwen3.6 35B A3B is the fastest at about 151.5 tokens/s. The full list is in the table above.

How many tokens per second does the GeForce RTX 4090D 48GB get on LLMs?

Estimated from memory bandwidth at 4-bit and 8K context: 8B dense: about 125.5 tokens/s, 14B dense: about 76.3 tokens/s, 32B dense: about 35.6 tokens/s, and 70B dense: about 16.9 tokens/s. MoE models read only the active parameters for each token, so they run much faster. Real-world results are usually 70%–100% of the estimate.

What LLMs can dual GeForce RTX 4090D 48GB cards run?

With vLLM tensor parallelism at 4-bit and 32K context, two GeForce RTX 4090D 48GB cards cannot fit any model of 50B or larger released in 2026. At 4-bit with 32K context the minimum is 4 cards, which fits Qwen3.8 Flash Next and DeepSeek V4 Flash.

How much does a GeForce RTX 4090D 48GB cost?

As of Sep 29, 2026: Xianyu median $3,526 (9 listings). Figures are medians of live listings, refreshed every 6 hours.

GeForce RTX 4090D 48GB vs RTX 6000 Ada: which is better for LLMs?

The RTX 6000 Ada has 48GB of VRAM and 960 GB/s of bandwidth; the GeForce RTX 4090D 48GB has 48GB and 1,008 GB/s. For 70B dense models at 4-bit, the GeForce RTX 4090D 48GB is estimated at about 16.9 tokens/s and the RTX 6000 Ada at about 16.1 tokens/s.