Tesla T10 Price & Used Prices: Which Local LLMs Can It Run?
Price alert
We'll email you when the Xianyu (CN) price reaches your target.
The Tesla T10 has 16GB of GDDR6 memory with 403 GB/s of bandwidth. The median asking price is about $199 on Xianyu. At 4-bit quantization and 8K context it fits a dense model of up to about 23B parameters, at roughly 20 tokens/s.
Xianyu
Tesla T10 price history
Xianyu
Buy links may carry affiliate tags; they do not change the price you pay.
Full Tesla T10 specifications
- Memory16 GB · GDDR6
- Bandwidth403 GB/s
- FP16 / BF1680 TFLOPS
- INT8160 TOPS
- Bus width
- 256-bit
- Architecture
- Turing (TU102)
- Cores
- 3,584
- Tensor cores
- 448
- Ecosystem
- CUDA
- Interface
- PCIe 3.0 x16
- NVLink
- Not supported
- Power
- 150 W
- Power connector
- 1× 8-pin
- Slots
- 1
- Length
- 267 mm
- Type
- Data center
- Launch date
- Feb 4, 2020
What LLMs can the Tesla T10 run?
Models up to 50B parameters on a single Tesla T10, versions released since 2026 only
| Model | Size | Runs? | Speed t/s | Max context |
|---|---|---|---|---|
| 3.8 GB | Runs | 72.1 | 128K | |
| 4.1 GB | Runs | 66.3 | 256K | |
| 4.7 GB | Runs | 57.4 | 256K | |
| 10.5 GB | Runs | 100.3 | 256K | |
| 11.8 GB | Runs | 24.1 | 64K | |
| 9.8 GB | Runs | 28.8 | 93K | |
| 11.9 GB | Runs | 114.3 | 75K | |
| 11.8 GB | Runs | 23.2 | 41K | |
| 12.3 GB | Runs | 129.2 | 181K |
| Model | Size | Runs? | Speed t/s | Max context |
|---|---|---|---|---|
| 5 GB | Runs | 55.8 | 128K | |
| 5.7 GB | Runs | 48.8 | 256K | |
| 7.1 GB | Runs | 39.2 | 256K | |
| 16.9 GB | Needs RAM offload | 35 | 256K | |
| 16.8 GB | Too big | — | — | |
| 16.5 GB | Too big | — | — | |
| 18.3 GB | Needs RAM offload | 35 | 198K | |
| 18.3 GB | Too big | — | — | |
| 22.1 GB | Needs RAM offload | 47.6 | 256K |
| Model | Size | Runs? | Speed t/s | Max context |
|---|---|---|---|---|
| 8.2 GB | Runs | 34.8 | 128K | |
| 9.5 GB | Runs | 30 | 196K | |
| 12.7 GB | Runs | 22.5 | 183K | |
| 26.9 GB | Needs RAM offload | 23.4 | 256K | |
| 28.6 GB | Too big | — | — | |
| 29 GB | Too big | — | — | |
| 31.8 GB | Needs RAM offload | 21.7 | 198K | |
| 32.6 GB | Too big | — | — | |
| 36.9 GB | Needs RAM offload | 31.1 | 256K |
RunsFits fully in VRAM; speed is estimated from VRAM bandwidth.
Needs RAM offloadMoE only: expert weights sit in system RAM, so speed is estimated from 70 GB/s RAM bandwidth.
Too bigWon’t fit on one card at this quantization, even with RAM offload.
Dual, 4x and 8x Tesla T10 for LLMs: which models fit?
Cards needed for 50B+ models released since 2026, with vLLM tensor parallelism at 32K context
| Model | Size | Cards | Total price | Total VRAM | Speed t/s | Total power |
|---|---|---|---|---|---|---|
| 78.9 GB | 8 cards | $1,589 | 128 GB | 69.4 | 1,200 W | |
| 96.8 GB | 8 cards | $1,589 | 128 GB | 51.9 | 1,200 W | |
| 108.7 GB | 8 cards | $1,589 | 128 GB | 39.6 | 1,200 W |
| Model | Size | Cards | Total price | Total VRAM | Speed t/s | Total power |
|---|---|---|---|---|---|---|
| 111.3 GB | 8 cards | $1,589 | 128 GB | 55.2 | 1,200 W |
At this quantization, even 8 × Tesla T10 cannot hold any model over 50B.
Card counts are 1 / 2 / 4 / 8, each using 90% of its VRAM. Speed is based on one card’s bandwidth; price covers GPUs only. No NVLink; cards talk over PCIe. Whether it runs also depends on vLLM support.
Other markets
Buy links may carry affiliate tags; they do not change the price you pay.
FAQ
How large a model can the Tesla T10 run?
A single Tesla T10 has 16GB of GDDR6 VRAM. At 4-bit quantization and 8K context it fits a dense model of up to about 23B parameters, or about 13B at 8-bit. MoE models can offload experts to system RAM to run larger ones, such as Qwen3.6 35B A3B (36B).
Can the Tesla T10 run a 70B model?
Not on a single card. At 4-bit with an 8K context a 70B dense model needs about 43GB, more than the Tesla T10's 16GB. It runs if you split the layers across 3 cards with llama.cpp (48GB total).
What are the best LLMs to run on the Tesla T10?
Among models released in 2026, the largest dense model that fits entirely in VRAM at 4-bit is Gemma 4 12B, at an estimated 39.2 tokens/s. The full list is in the table above.
How many tokens per second does the Tesla T10 get on LLMs?
Estimated from memory bandwidth at 4-bit and 8K context: 8B dense: about 53.4 tokens/s and 14B dense: about 31.7 tokens/s. MoE models read only the active parameters for each token, so they run much faster. Real-world results are usually 70%–100% of the estimate.
What LLMs can dual Tesla T10 cards run?
With vLLM tensor parallelism at 4-bit and 32K context, two Tesla T10 cards cannot fit any model of 50B or larger released in 2026. At 4-bit with 32K context the minimum is 8 cards, which fits Qwen3.8 Flash Next.
How much does a Tesla T10 cost?
As of Sep 29, 2026: eBay median $495 (5 listings) and Xianyu median $199 (9 listings). Figures are medians of live listings, refreshed every 6 hours.
Is the Tesla T10 price going up or down?
As of Sep 29, 2026, the Xianyu median is up 2.7% over 7 days, up 2.7% over 30 days.
Tesla T10 vs Tesla P40: which is better for LLMs?
The Tesla P40 has 24GB of VRAM and 347 GB/s of bandwidth; the Tesla T10 has 16GB and 403 GB/s. At the Xianyu median, the Tesla P40 costs about $179, 10.1% less than the Tesla T10 ($199). For 14B dense models at 4-bit, the Tesla T10 is estimated at about 31.7 tokens/s and the Tesla P40 at about 27.4 tokens/s.