Tesla P40 Price & Used Prices: Which Local LLMs Can It Run?
Price alert
We'll email you when the eBay (US) price reaches your target.
The Tesla P40 has 24GB of GDDR5 memory with 347 GB/s of bandwidth. The median asking price is about $397 on eBay. At 4-bit quantization and 8K context it fits a dense model of up to about 37B parameters, at roughly 11 tokens/s.
eBay
Tesla P40 price history
eBay
Buy links may carry affiliate tags; they do not change the price you pay.
Full Tesla P40 specifications
- Memory24 GB · GDDR5
- Bandwidth347 GB/s
- INT847 TOPS
- Bus width
- 384-bit
- Architecture
- Pascal (GP102)
- Cores
- 3,840
- Ecosystem
- CUDA
- Interface
- PCIe 3.0 x16
- NVLink
- Not supported
- Power
- 250 W
- Power connector
- 8-pin CPU (EPS-12V)
- Slots
- 2
- Length
- 267 mm
- Type
- Data center
- Launch date
- Sep 13, 2016
What LLMs can the Tesla P40 run?
Models up to 50B parameters on a single Tesla P40, versions released since 2026 only
| Model | Size | Runs? | Speed t/s | Max context |
|---|---|---|---|---|
| 3.8 GB | Runs | 62.5 | 128K | |
| 4.1 GB | Runs | 57.5 | 256K | |
| 4.7 GB | Runs | 49.7 | 256K | |
| 10.5 GB | Runs | 91.5 | 256K | |
| 11.8 GB | Runs | 20.8 | 192K | |
| 9.8 GB | Runs | 24.8 | 221K | |
| 11.9 GB | Runs | 105.1 | 198K | |
| 11.8 GB | Runs | 20 | 143K | |
| 12.3 GB | Runs | 119.8 | 256K |
| Model | Size | Runs? | Speed t/s | Max context |
|---|---|---|---|---|
| 5 GB | Runs | 48.3 | 128K | |
| 5.7 GB | Runs | 42.3 | 256K | |
| 7.1 GB | Runs | 33.9 | 256K | |
| 16.9 GB | Runs | 68 | 256K | |
| 16.8 GB | Runs | 14.8 | 117K | |
| 16.5 GB | Runs | 15 | 122K | |
| 18.3 GB | Runs | 83.2 | 115K | |
| 18.3 GB | Runs | 13.3 | 66K | |
| 22.1 GB | Runs | 86.6 | 123K |
| Model | Size | Runs? | Speed t/s | Max context |
|---|---|---|---|---|
| 8.2 GB | Runs | 30.1 | 128K | |
| 9.5 GB | Runs | 25.9 | 256K | |
| 12.7 GB | Runs | 19.4 | 256K | |
| 26.9 GB | Needs RAM offload | 22.6 | 256K | |
| 28.6 GB | Too big | — | — | |
| 29 GB | Too big | — | — | |
| 31.8 GB | Needs RAM offload | 21.3 | 198K | |
| 32.6 GB | Too big | — | — | |
| 36.9 GB | Needs RAM offload | 30 | 256K |
RunsFits fully in VRAM; speed is estimated from VRAM bandwidth.
Needs RAM offloadMoE only: expert weights sit in system RAM, so speed is estimated from 70 GB/s RAM bandwidth.
Too bigWon’t fit on one card at this quantization, even with RAM offload.
Dual, 4x and 8x Tesla P40 for LLMs: which models fit?
Cards needed for 50B+ models released since 2026, with vLLM tensor parallelism at 32K context
| Model | Size | Cards | Total price | Total VRAM | Speed t/s | Total power |
|---|---|---|---|---|---|---|
| 78.9 GB | 4 cards | $1,588 | 96 GB | 62.1 | 1,000 W | |
| 96.8 GB | 8 cards | $3,176 | 192 GB | 46 | 2,000 W | |
| 107.2 GB | 8 cards | $3,176 | 192 GB | 18.3 | 2,000 W | |
| 108.7 GB | 8 cards | $3,176 | 192 GB | 34.9 | 2,000 W |
| Model | Size | Cards | Total price | Total VRAM | Speed t/s | Total power |
|---|---|---|---|---|---|---|
| 111.3 GB | 8 cards | $3,176 | 192 GB | 49 | 2,000 W | |
| 155.1 GB | 8 cards | $3,176 | 192 GB | 31.1 | 2,000 W |
| Model | Size | Cards | Total price | Total VRAM | Speed t/s | Total power |
|---|---|---|---|---|---|---|
| 161.9 GB | 8 cards | $3,176 | 192 GB | 29.9 | 2,000 W |
Card counts are 1 / 2 / 4 / 8, each using 90% of its VRAM. Speed is based on one card’s bandwidth; price covers GPUs only. No NVLink; cards talk over PCIe. Whether it runs also depends on vLLM support.
Tesla P40 rental price per hour
On Compshare, the Tesla P40 rents for $0.06 per hour on demand, about $14/month at 8 hours a day. Buying one (eBay median $397) pays off in about 23.6 years, 11.8 years or 3.9 years at 4, 8 or 24 hours a day.
| Option | Price | 4 h/day | 8 h/day | 24 h/day | View |
|---|---|---|---|---|---|
| $0.06/hr | $7/mo | $14/mo | $41/mo | Rent | |
| $397 | Payback ~23.6 yrsPower $5/mo | Payback ~11.8 yrsPower $11/mo | Payback ~3.9 yrsPower $33/mo | Buy |
Break-even uses the US residential average ($0.18/kWh) and 250 W rated power, excluding resale value and the rest of the build.
Other markets
Buy links may carry affiliate tags; they do not change the price you pay.
FAQ
How large a model can the Tesla P40 run?
A single Tesla P40 has 24GB of GDDR5 VRAM. At 4-bit quantization and 8K context it fits a dense model of up to about 37B parameters, or about 20B at 8-bit.
Can the Tesla P40 run a 70B model?
Not on a single card. At 4-bit with an 8K context a 70B dense model needs about 43GB, more than the Tesla P40's 24GB. It runs if you split the layers across 2 cards with llama.cpp (48GB total).
What are the best LLMs to run on the Tesla P40?
Among models released in 2026, the largest dense model that fits entirely in VRAM at 4-bit is Gemma 4 31B, at an estimated 13.3 tokens/s. Among MoE models, Qwen3.6 35B A3B is the fastest at about 86.6 tokens/s. The full list is in the table above.
How many tokens per second does the Tesla P40 get on LLMs?
Estimated from memory bandwidth at 4-bit and 8K context: 8B dense: about 46.2 tokens/s, 14B dense: about 27.4 tokens/s, and 32B dense: about 12.5 tokens/s. MoE models read only the active parameters for each token, so they run much faster. Real-world results are usually 70%–100% of the estimate.
What LLMs can dual Tesla P40 cards run?
With vLLM tensor parallelism at 4-bit and 32K context, two Tesla P40 cards cannot fit any model of 50B or larger released in 2026. At 4-bit with 32K context the minimum is 8 cards, which fits Qwen3.8 Flash Next and DeepSeek V4 Flash.
How much does a Tesla P40 cost?
As of Sep 29, 2026: eBay median $397 (16 listings) and Xianyu median $179 (13 listings). Figures are medians of live listings, refreshed every 6 hours.
Is the Tesla P40 price going up or down?
As of Sep 29, 2026, the eBay median is up 2.6% over 7 days, up 2.6% over 30 days.
How much does it cost to rent a Tesla P40 per hour?
As of Sep 29, 2026: Compshare at $0.06 per hour. All are single-GPU on-demand rates, excluding storage and data transfer.
Is it cheaper to buy or rent a Tesla P40 at 8 hours a day?
Taking the eBay median of $397 and electricity at $0.18 per kWh, at 8 hours a day buying pays for itself in about 11.8 years (versus Compshare at $0.06 per hour). If you will use it for longer than 11.8 years, buy; otherwise, rent.
Tesla P40 vs GeForce RTX 3060 12GB: which is better for LLMs?
The GeForce RTX 3060 12GB has 12GB of VRAM and 360 GB/s of bandwidth; the Tesla P40 has 24GB and 347 GB/s. At the eBay median, the GeForce RTX 3060 12GB costs about $420, 5.8% more than the Tesla P40 ($397). For 14B dense models at 4-bit, the Tesla P40 is estimated at about 27.4 tokens/s and the GeForce RTX 3060 12GB at about 28.4 tokens/s.