Tesla V100 16GB: 16GB VRAM — What Local LLMs Can It Run?
Check prices
The Tesla V100 16GB has 16GB of HBM2 memory with 900 GB/s of bandwidth. No used-price data yet. At Q4 quantization it runs models up to 14B parameters at roughly 69 tokens/s.
Full Tesla V100 16GB specifications
What LLMs can the Tesla V100 16GB run?
One reference model per family (its main version)| Model | Q4_K_M | Q8_0 | FP16 | ||||||
|---|---|---|---|---|---|---|---|---|---|
| Fits | Speed t/s | Context | Fits | Speed t/s | Context | Fits | Speed t/s | Context | |
| DeepSeekDeepSeek R1 Qwen3 8B | ✓ | 109.9 | 32K | ✓ | 64.6 | 32K | ✗ | — | — |
| GemmaGemma 4 31B | ✗ | — | — | ✗ | — | — | ✗ | — | — |
| GLMGLM 4.7 Flash | Offload | 37.9 | — | Offload | 22.1 | — | Offload | 12.3 | — |
| gpt-ossgpt-oss 20B | ✓ | 132.4 | 32K | Offload | 16.8 | — | Offload | 9.2 | — |
| KimiKimi Linear 48B A3B | Offload | 44 | — | Offload | 25.8 | — | Offload | 14.4 | — |
| LlamaLlama 3.1 8B | ✓ | 112.9 | 32K | ✓ | 66.2 | 32K | ✗ | — | — |
| MistralMistral 7B v0.3 | ✓ | 122.7 | 32K | ✓ | 72.5 | 32K | ✗ | — | — |
| QwenQwen3 8B | ✓ | 109.9 | 32K | ✓ | 64.6 | 32K | ✗ | — | — |
Fits when weights + KV cache + runtime overhead is at or below usable memory, single card. "Offload" on MoE models means it runs with expert weights in system RAM (only attention layers and KV cache stay in VRAM); that speed assumes 70 GB/s RAM bandwidth. The context column is the largest tier that fits.
Tesla V100 16GB used price history
Not enough history to draw a chart yet — price tracking just started.
Prices by platform
Buy links may carry affiliate tags; they do not change the price you pay.
FAQ
How much VRAM does the Tesla V100 16GB have?
The Tesla V100 16GB has 16GB of HBM2 memory with 900 GB/s of bandwidth.
Can the Tesla V100 16GB run a 70B model?
Not on a single card. At Q4_K_M with an 8K context a 70B dense model needs about 43GB, more than the Tesla V100 16GB's 16GB.
How much does a used Tesla V100 16GB cost?
No used-price data yet; it will appear here once the feed is live.