RTX PRO 5500 Price & Used Prices: Which Local LLMs Can It Run?
The RTX PRO 5500 has 84GB of GDDR7 ECC memory with 1,398 GB/s of bandwidth. No price data yet. At 4-bit quantization and 8K context it fits a dense model of up to about 143B parameters, at roughly 12 tokens/s.
eBay
RTX PRO 5500 price history
eBay
Buy links may carry affiliate tags; they do not change the price you pay.
Full RTX PRO 5500 specifications
- Memory84 GB · GDDR7 ECC
- Bandwidth1,398 GB/s
- Bus width
- 448-bit
- Architecture
- Blackwell (GB202)
- Cores
- 21,760
- Tensor cores
- 680
- Ecosystem
- CUDA
- Interface
- PCIe 5.0 x16
- NVLink
- Not supported
- Power
- 600 W
- Slots
- 2
- Length
- 282 mm
- Type
- Workstation
- Launch date
- Sep 14, 2026
What LLMs can the RTX PRO 5500 run?
Models up to 50B parameters on a single RTX PRO 5500, versions released since 2026 only
| Model | Size | Runs? | Speed t/s | Max context |
|---|---|---|---|---|
| 3.8 GB | Runs | 218.7 | 128K | |
| 4.1 GB | Runs | 203.3 | 256K | |
| 4.7 GB | Runs | 178.6 | 256K | |
| 10.5 GB | Runs | 174.8 | 256K | |
| 11.8 GB | Runs | 79.7 | 256K | |
| 9.8 GB | Runs | 94.4 | 256K | |
| 11.9 GB | Runs | 186.2 | 198K | |
| 11.8 GB | Runs | 76.9 | 256K | |
| 12.3 GB | Runs | 196.9 | 256K |
| Model | Size | Runs? | Speed t/s | Max context |
|---|---|---|---|---|
| 5 GB | Runs | 174.2 | 128K | |
| 5.7 GB | Runs | 154.4 | 256K | |
| 7.1 GB | Runs | 126.1 | 256K | |
| 16.9 GB | Runs | 150.2 | 256K | |
| 16.8 GB | Runs | 57.5 | 256K | |
| 16.5 GB | Runs | 58.4 | 256K | |
| 18.3 GB | Runs | 166.9 | 198K | |
| 18.3 GB | Runs | 51.7 | 256K | |
| 22.1 GB | Runs | 170.2 | 256K |
| Model | Size | Runs? | Speed t/s | Max context |
|---|---|---|---|---|
| 8.2 GB | Runs | 112.9 | 128K | |
| 9.5 GB | Runs | 98.3 | 256K | |
| 12.7 GB | Runs | 74.7 | 256K | |
| 26.9 GB | Runs | 123 | 256K | |
| 28.6 GB | Runs | 34.6 | 256K | |
| 29 GB | Runs | 34.2 | 256K | |
| 31.8 GB | Runs | 136.9 | 198K | |
| 32.6 GB | Runs | 30.1 | 256K | |
| 36.9 GB | Runs | 141.3 | 256K |
RunsFits fully in VRAM; speed is estimated from VRAM bandwidth.
Needs RAM offloadMoE only: expert weights sit in system RAM, so speed is estimated from 70 GB/s RAM bandwidth.
Too bigWon’t fit on one card at this quantization, even with RAM offload.
Dual, 4x and 8x RTX PRO 5500 for LLMs: which models fit?
Cards needed for 50B+ models released since 2026, with vLLM tensor parallelism at 32K context
| Model | Size | Cards | Total price | Total VRAM | Speed t/s | Total power |
|---|---|---|---|---|---|---|
| 78.9 GB | 1 card | No price | 84 GB | 142.8 | 600 W | |
| 96.8 GB | 2 cards | No price | 168 GB | 119.1 | 1,200 W | |
| 107.2 GB | 2 cards | No price | 168 GB | 60.5 | 1,200 W | |
| 108.7 GB | 2 cards | No price | 168 GB | 98.8 | 1,200 W | |
| 253.9 GB | 4 cards | No price | 336 GB | 53.5 | 2,400 W | |
| 286.1 GBest. | 4 cards | No price | 336 GB | 40.9 | 2,400 W | |
| 339.5 GB | 8 cards | No price | 672 GB | 63.8 | 4,800 W | |
| 594.6 GBest. | 8 cards | No price | 672 GB | 43.2 | 4,800 W |
| Model | Size | Cards | Total price | Total VRAM | Speed t/s | Total power |
|---|---|---|---|---|---|---|
| 111.3 GB | 2 cards | No price | 168 GB | 123.8 | 1,200 W | |
| 155.1 GB | 2 cards | No price | 168 GB | 90.9 | 1,200 W | |
| 182.2 GB | 4 cards | No price | 336 GB | 46 | 2,400 W | |
| 199.7 GB | 4 cards | No price | 336 GB | 66.2 | 2,400 W | |
| 467.3 GB | 8 cards | No price | 672 GB | 33.5 | 4,800 W | |
| 467.3 GB | 8 cards | No price | 672 GB | 29.1 | 4,800 W | |
| 510.3 GB | 8 cards | No price | 672 GB | 57.7 | 4,800 W | |
| 583.7 GB | 8 cards | No price | 672 GB | 43 | 4,800 W |
| Model | Size | Cards | Total price | Total VRAM | Speed t/s | Total power |
|---|---|---|---|---|---|---|
| 161.9 GB | 4 cards | No price | 336 GB | 88.5 | 2,400 W | |
| 188.2 GB | 4 cards | No price | 336 GB | 94.1 | 2,400 W | |
| 317.7 GB | 8 cards | No price | 672 GB | 32.1 | 4,800 W | |
| 341 GB | 8 cards | No price | 672 GB | 43.8 | 4,800 W | |
| 594.5 GB | 8 cards | No price | 672 GB | 42.4 | 4,800 W |
Card counts are 1 / 2 / 4 / 8, each using 90% of its VRAM. Speed is based on one card’s bandwidth; price covers GPUs only. Rows marked “Engram in RAM” need about 203 GB of extra system RAM. No NVLink; cards talk over PCIe. Whether it runs also depends on vLLM support.
Other markets
Buy links may carry affiliate tags; they do not change the price you pay.
FAQ
How large a model can the RTX PRO 5500 run?
A single RTX PRO 5500 has 84GB of GDDR7 ECC VRAM. At 4-bit quantization and 8K context it fits a dense model of up to about 143B parameters, or about 77B at 8-bit.
Can the RTX PRO 5500 run a 70B model?
Yes. At 4-bit with an 8K context a 70B dense model needs about 43GB, which fits in the RTX PRO 5500's 84GB.
What are the best LLMs to run on the RTX PRO 5500?
Among models released in 2026, the largest dense model that fits entirely in VRAM at 4-bit is Gemma 4 31B, at an estimated 51.7 tokens/s. Among MoE models, Qwen3.6 35B A3B is the fastest at about 170.2 tokens/s. The full list is in the table above.
How many tokens per second does the RTX PRO 5500 get on LLMs?
Estimated from memory bandwidth at 4-bit and 8K context: 8B dense: about 167.5 tokens/s, 14B dense: about 103.4 tokens/s, 32B dense: about 48.9 tokens/s, and 70B dense: about 23.3 tokens/s. MoE models read only the active parameters for each token, so they run much faster. Real-world results are usually 70%–100% of the estimate.
What LLMs can dual RTX PRO 5500 cards run?
Two RTX PRO 5500 cards (168GB total) with vLLM tensor parallelism at 4-bit and 32K context can fit Qwen3.8 Flash Next and DeepSeek V4 Flash. To run Kimi K2.7 Code at 4-bit with 32K context, you need 8 cards.
How much does a RTX PRO 5500 cost?
No price data yet; it will appear here once the feed is live.