A100 80GB PCIe Price & Used Prices: Which Local LLMs Can It Run?
The A100 80GB PCIe has 80GB of HBM2e memory with 1,935 GB/s of bandwidth. No price data yet. At 4-bit quantization and 8K context it fits a dense model of up to about 136B parameters, at roughly 17 tokens/s.
Amazon.jp
A100 80GB PCIe price history
Amazon.jp
Buy links may carry affiliate tags; they do not change the price you pay.
Full A100 80GB PCIe specifications
- Memory80 GB · HBM2e
- Bandwidth1,935 GB/s
- FP16 / BF16312 TFLOPS
- INT8624 TOPS
- Bus width
- 5,120-bit
- Architecture
- Ampere (GA100)
- Cores
- 6,912
- Tensor cores
- 432
- Ecosystem
- CUDA
- Interface
- PCIe 4.0 x16
- NVLink
- 600 GB/s
- Power
- 300 W
- Power connector
- 8-pin EPS
- Slots
- 2
- Length
- 267 mm
- Type
- Data center
- Launch date
- Jun 28, 2021
What LLMs can the A100 80GB PCIe run?
Models up to 50B parameters on a single A100 80GB PCIe, versions released since 2026 only
| Model | Size | Runs? | Speed t/s | Max context |
|---|---|---|---|---|
| 3.8 GB | Runs | 283.7 | 128K | |
| 4.1 GB | Runs | 264.9 | 256K | |
| 4.7 GB | Runs | 234.4 | 256K | |
| 10.5 GB | Runs | 190.7 | 256K | |
| 11.8 GB | Runs | 107.7 | 256K | |
| 9.8 GB | Runs | 127 | 256K | |
| 11.9 GB | Runs | 200.4 | 198K | |
| 11.8 GB | Runs | 104 | 256K | |
| 12.3 GB | Runs | 209.2 | 256K |
| Model | Size | Runs? | Speed t/s | Max context |
|---|---|---|---|---|
| 5 GB | Runs | 228.9 | 128K | |
| 5.7 GB | Runs | 204.1 | 256K | |
| 7.1 GB | Runs | 168 | 256K | |
| 16.9 GB | Runs | 168.9 | 256K | |
| 16.8 GB | Runs | 78.1 | 256K | |
| 16.5 GB | Runs | 79.5 | 256K | |
| 18.3 GB | Runs | 183.9 | 198K | |
| 18.3 GB | Runs | 70.5 | 256K | |
| 22.1 GB | Runs | 186.8 | 256K |
| Model | Size | Runs? | Speed t/s | Max context |
|---|---|---|---|---|
| 8.2 GB | Runs | 151.1 | 128K | |
| 9.5 GB | Runs | 132.1 | 256K | |
| 12.7 GB | Runs | 101.1 | 256K | |
| 26.9 GB | Runs | 143.2 | 256K | |
| 28.6 GB | Runs | 47.4 | 256K | |
| 29 GB | Runs | 46.8 | 256K | |
| 31.8 GB | Runs | 156.5 | 198K | |
| 32.6 GB | Runs | 41.3 | 256K | |
| 36.9 GB | Runs | 160.7 | 256K |
RunsFits fully in VRAM; speed is estimated from VRAM bandwidth.
Needs RAM offloadMoE only: expert weights sit in system RAM, so speed is estimated from 70 GB/s RAM bandwidth.
Too bigWon’t fit on one card at this quantization, even with RAM offload.
Dual, 4x and 8x A100 80GB PCIe for LLMs: which models fit?
Cards needed for 50B+ models released since 2026, with vLLM tensor parallelism at 32K context
| Model | Size | Cards | Total price | Total VRAM | Speed t/s | Total power |
|---|---|---|---|---|---|---|
| 78.9 GB | 2 cards | No price | 160 GB | 162.1 | 600 W | |
| 96.8 GB | 2 cards | No price | 160 GB | 139.3 | 600 W | |
| 107.2 GB | 2 cards | No price | 160 GB | 76.6 | 600 W | |
| 108.7 GB | 2 cards | No price | 160 GB | 118.7 | 600 W | |
| 253.9 GB | 4 cards | No price | 320 GB | 68.4 | 1,200 W | |
| 286.1 GBest. | 8 cards | No price | 640 GB | 53.3 | 2,400 W | |
| 339.5 GB | 8 cards | No price | 640 GB | 80.4 | 2,400 W |
| Model | Size | Cards | Total price | Total VRAM | Speed t/s | Total power |
|---|---|---|---|---|---|---|
| 111.3 GB | 2 cards | No price | 160 GB | 144 | 600 W | |
| 155.1 GB | 4 cards | No price | 320 GB | 110.4 | 1,200 W | |
| 182.2 GB | 4 cards | No price | 320 GB | 59.5 | 1,200 W | |
| 199.7 GB | 4 cards | No price | 320 GB | 83.2 | 1,200 W | |
| 467.3 GB | 8 cards | No price | 640 GB | 44.1 | 2,400 W | |
| 467.3 GB | 8 cards | No price | 640 GB | 38.6 | 2,400 W | |
| 510.3 GB | 8 cardsEngram in RAM | No price | 640 GB | 73.3 | 2,400 W | |
| 583.7 GB | 8 cards | No price | 640 GB | 55.8 | 2,400 W |
| Model | Size | Cards | Total price | Total VRAM | Speed t/s | Total power |
|---|---|---|---|---|---|---|
| 161.9 GB | 4 cards | No price | 320 GB | 107.8 | 1,200 W | |
| 188.2 GB | 4 cards | No price | 320 GB | 113.8 | 1,200 W | |
| 317.7 GB | 8 cards | No price | 640 GB | 42.4 | 2,400 W | |
| 341 GB | 8 cards | No price | 640 GB | 56.8 | 2,400 W | |
| 594.5 GB | 8 cards | No price | 640 GB | 55 | 2,400 W |
Card counts are 1 / 2 / 4 / 8, each using 90% of its VRAM. Speed is based on one card’s bandwidth; price covers GPUs only. Rows marked “Engram in RAM” need about 203 GB of extra system RAM. Supports NVLink (600 GB/s). Whether it runs also depends on vLLM support.
A100 80GB PCIe rental price per hour
On-demand rental of the A100 80GB PCIe starts at $1 per hour (Vast.ai), about $244/month at 8 hours a day. Buying one (eBay median $17,637) pays off in about 12.7 years, 6.4 years or 2.1 years at 4, 8 or 24 hours a day.
| Option | Price | 4 h/day | 8 h/day | 24 h/day | View |
|---|---|---|---|---|---|
| $1/hr | $122/mo | $244/mo | $732/mo | Rent | |
| $1.19/hr | $145/mo | $290/mo | $869/mo | Rent | |
| $17,637 | Payback ~12.7 yrsPower $7/mo | Payback ~6.4 yrsPower $13/mo | Payback ~2.1 yrsPower $39/mo | Buy |
Break-even uses the US residential average ($0.18/kWh) and 300 W rated power, excluding resale value and the rest of the build.
Other markets
Buy links may carry affiliate tags; they do not change the price you pay.
FAQ
How large a model can the A100 80GB PCIe run?
A single A100 80GB PCIe has 80GB of HBM2e VRAM. At 4-bit quantization and 8K context it fits a dense model of up to about 136B parameters, or about 73B at 8-bit.
Can the A100 80GB PCIe run a 70B model?
Yes. At 4-bit with an 8K context a 70B dense model needs about 43GB, which fits in the A100 80GB PCIe's 80GB.
What are the best LLMs to run on the A100 80GB PCIe?
Among models released in 2026, the largest dense model that fits entirely in VRAM at 4-bit is Gemma 4 31B, at an estimated 70.5 tokens/s. Among MoE models, Qwen3.6 35B A3B is the fastest at about 186.8 tokens/s. The full list is in the table above.
How many tokens per second does the A100 80GB PCIe get on LLMs?
Estimated from memory bandwidth at 4-bit and 8K context: 8B dense: about 220.5 tokens/s, 14B dense: about 138.7 tokens/s, 32B dense: about 66.6 tokens/s, and 70B dense: about 32 tokens/s. MoE models read only the active parameters for each token, so they run much faster. Real-world results are usually 70%–100% of the estimate.
What LLMs can dual A100 80GB PCIe cards run?
Two A100 80GB PCIe cards (160GB total) with vLLM tensor parallelism at 4-bit and 32K context can fit Qwen3.8 Flash Next. To run Kimi K2.7 Code at 4-bit with 32K context, you need 8 cards.
How much does a A100 80GB PCIe cost?
As of Sep 29, 2026: eBay median $17,637 (8 listings) and Xianyu median $17,842 (4 listings). Figures are medians of live listings, refreshed every 6 hours.
How much does it cost to rent a A100 80GB PCIe per hour?
As of Sep 29, 2026: Vast.ai at $1 per hour and RunPod at $1.19 per hour. All are single-GPU on-demand rates, excluding storage and data transfer.
Is it cheaper to buy or rent a A100 80GB PCIe at 8 hours a day?
Taking the eBay median of $17,637 and electricity at $0.18 per kWh, at 8 hours a day buying pays for itself in about 6.4 years (versus Vast.ai at $1 per hour). If you will use it for longer than 6.4 years, buy; otherwise, rent.
A100 80GB PCIe vs H100 PCIe: which is better for LLMs?
The H100 PCIe has 80GB of VRAM and 2,000 GB/s of bandwidth; the A100 80GB PCIe has 80GB and 1,935 GB/s. For 70B dense models at 4-bit, the A100 80GB PCIe is estimated at about 32 tokens/s and the H100 PCIe at about 33.1 tokens/s.