GPUs tracked56Price history27 daysPrice points (24h)34,228 pointsBiggest 24h dropA100 80GB PCIe-20.0%Biggest 24h riseRTX PRO 5000 48GB+17.5%Updated
$ USD
English

L40S Price & Used Prices: Which Local LLMs Can It Run?

The L40S has 48GB of GDDR6 memory with 864 GB/s of bandwidth. No price data yet. At 4-bit quantization and 8K context it fits a dense model of up to about 79B parameters, at roughly 13 tokens/s.

L40S price history

Price updated —

Not enough data to show a trend

The median price trend will appear once enough valid quotes are available.

Buy links may carry affiliate tags; they do not change the price you pay.

Full L40S specifications

  • Memory48 GB · GDDR6
  • Bandwidth864 GB/s
  • FP16 / BF16362.1 TFLOPS
  • FP8 / INT8733 / 733 TOPS
Bus width
384-bit
Architecture
Ada Lovelace (AD102)
Cores
18,176
Tensor cores
568
Ecosystem
CUDA
Interface
PCIe 4.0 x16
NVLink
Not supported
Power
350 W
Power connector
16-pin
Slots
2
Length
267 mm
Type
Data center
Launch date
Aug 8, 2023

What LLMs can the L40S run?

Models up to 50B parameters on a single L40S, versions released since 2026 only

ModelSizeRuns?Speed t/sMax context
Gemma 4 E4B8B5 GBRuns113.7128K
Qwen3.5 9B9.7B5.7 GBRuns100.2256K
Gemma 4 12B12B7.1 GBRuns81256K
Gemma 4 26B A4B25.2B (3.8B active)16.9 GBRuns120.4256K
Qwen3.6 27B27.8B16.8 GBRuns36.1256K
Qwen3.8 27B27.8B16.5 GBRuns36.8256K
GLM 4.7 Flash30B (3B active)18.3 GBRuns138.4198K
Gemma 4 31B30.7B18.3 GBRuns32.5256K
Qwen3.6 35B A3B36B (3B active)22.1 GBRuns142.2256K

RunsFits fully in VRAM; speed is estimated from VRAM bandwidth.

Needs RAM offloadMoE only: expert weights sit in system RAM, so speed is estimated from 70 GB/s RAM bandwidth.

Too bigWon’t fit on one card at this quantization, even with RAM offload.

Dual, 4x and 8x L40S for LLMs: which models fit?

Cards needed for 50B+ models released since 2026, with vLLM tensor parallelism at 32K context

ModelSizeCardsTotal priceTotal VRAMSpeed t/sTotal power
Qwen3.8 Flash Next176B (6B active)111.3 GB4 cardsNo price192 GB94.41,400 W
DeepSeek V4 Flash284B (13B active)155.1 GB4 cardsNo price192 GB65.31,400 W
Hy3295B (21B active)182.2 GB8 cardsNo price384 GB30.62,800 W
GLM 5.3 Flash320B (18B active)199.7 GB8 cardsNo price384 GB45.52,800 W

Card counts are 1 / 2 / 4 / 8, each using 90% of its VRAM. Speed is based on one card’s bandwidth; price covers GPUs only. No NVLink; cards talk over PCIe. Whether it runs also depends on vLLM support.

L40S rental price per hour

On-demand rental of the L40S starts at $0.79 per hour (RunPod), about $192/month at 8 hours a day. Buying one (eBay median $8,974) pays off in about 8.5 years, 4.2 years or 1.4 years at 4, 8 or 24 hours a day.

OptionPrice4 h/day8 h/day24 h/dayView
RunPodRentUpdated 19 min. ago$0.79/hr$96/mo$192/mo$577/moRent
Vast.aiRentUpdated 19 min. ago$0.8/hr$98/mo$195/mo$585/moRent
eBayUpdated 4 hr. ago$8,974Payback ~8.5 yrsPower $8/moPayback ~4.2 yrsPower $15/moPayback ~1.4 yrsPower $46/moBuy

Break-even uses the US residential average ($0.18/kWh) and 350 W rated power, excluding resale value and the rest of the build.

Other markets

eBay · United States
$8,974
18 listings
View
Amazon.jp · Japan
N/A
No listings
View

Buy links may carry affiliate tags; they do not change the price you pay.

FAQ

How large a model can the L40S run?

A single L40S has 48GB of GDDR6 VRAM. At 4-bit quantization and 8K context it fits a dense model of up to about 79B parameters, or about 43B at 8-bit.

Can the L40S run a 70B model?

Yes. At 4-bit with an 8K context a 70B dense model needs about 43GB, which fits in the L40S's 48GB.

What are the best LLMs to run on the L40S?

Among models released in 2026, the largest dense model that fits entirely in VRAM at 4-bit is Gemma 4 31B, at an estimated 32.5 tokens/s. Among MoE models, Qwen3.6 35B A3B is the fastest at about 142.2 tokens/s. The full list is in the table above.

How many tokens per second does the L40S get on LLMs?

Estimated from memory bandwidth at 4-bit and 8K context: 8B dense: about 109.1 tokens/s, 14B dense: about 66 tokens/s, 32B dense: about 30.7 tokens/s, and 70B dense: about 14.5 tokens/s. MoE models read only the active parameters for each token, so they run much faster. Real-world results are usually 70%–100% of the estimate.

What LLMs can dual L40S cards run?

With vLLM tensor parallelism at 4-bit and 32K context, two L40S cards cannot fit any model of 50B or larger released in 2026. At 4-bit with 32K context the minimum is 4 cards, which fits Qwen3.8 Flash Next and DeepSeek V4 Flash.

How much does a L40S cost?

As of Sep 29, 2026: eBay median $8,974 (18 listings). Figures are medians of live listings, refreshed every 6 hours.

How much does it cost to rent a L40S per hour?

As of Sep 29, 2026: RunPod at $0.79 per hour and Vast.ai at $0.8 per hour. All are single-GPU on-demand rates, excluding storage and data transfer.

Is it cheaper to buy or rent a L40S at 8 hours a day?

Taking the eBay median of $8,974 and electricity at $0.18 per kWh, at 8 hours a day buying pays for itself in about 4.2 years (versus RunPod at $0.79 per hour). If you will use it for longer than 4.2 years, buy; otherwise, rent.

L40S vs RTX 6000 Ada: which is better for LLMs?

The RTX 6000 Ada has 48GB of VRAM and 960 GB/s of bandwidth; the L40S has 48GB and 864 GB/s. For 70B dense models at 4-bit, the L40S is estimated at about 14.5 tokens/s and the RTX 6000 Ada at about 16.1 tokens/s.