GPUs tracked56Price history27 daysPrice points (24h)34,228 pointsBiggest 24h dropA100 80GB PCIe-20.0%Biggest 24h riseRTX PRO 5000 48GB+17.5%Updated
$ USD
English

H100 PCIe Price & Used Prices: Which Local LLMs Can It Run?

The H100 PCIe has 80GB of HBM2e memory with 2,000 GB/s of bandwidth. No price data yet. At 4-bit quantization and 8K context it fits a dense model of up to about 136B parameters, at roughly 18 tokens/s.

H100 PCIe price history

Price updated —

Not enough data to show a trend

The median price trend will appear once enough valid quotes are available.

Buy links may carry affiliate tags; they do not change the price you pay.

Full H100 PCIe specifications

  • Memory80 GB · HBM2e
  • Bandwidth2,000 GB/s
  • FP16 / BF16756 TFLOPS
  • FP8 / INT81,513 / 1,513 TOPS
Bus width
5,120-bit
Architecture
Hopper (GH100)
Cores
14,592
Tensor cores
456
Ecosystem
CUDA
Interface
PCIe 5.0 x16
NVLink
600 GB/s
Power
350 W
Power connector
16-pin
Slots
2
Length
268 mm
Type
Data center
Launch date
Mar 22, 2022

What LLMs can the H100 PCIe run?

Models up to 50B parameters on a single H100 PCIe, versions released since 2026 only

ModelSizeRuns?Speed t/sMax context
Gemma 4 E4B8B5 GBRuns235.1128K
Qwen3.5 9B9.7B5.7 GBRuns209.8256K
Gemma 4 12B12B7.1 GBRuns172.8256K
Gemma 4 26B A4B25.2B (3.8B active)16.9 GBRuns170.7256K
Qwen3.6 27B27.8B16.8 GBRuns80.6256K
Qwen3.8 27B27.8B16.5 GBRuns81.9256K
GLM 4.7 Flash30B (3B active)18.3 GBRuns185.4198K
Gemma 4 31B30.7B18.3 GBRuns72.7256K
Qwen3.6 35B A3B36B (3B active)22.1 GBRuns188.3256K

RunsFits fully in VRAM; speed is estimated from VRAM bandwidth.

Needs RAM offloadMoE only: expert weights sit in system RAM, so speed is estimated from 70 GB/s RAM bandwidth.

Too bigWon’t fit on one card at this quantization, even with RAM offload.

Dual, 4x and 8x H100 PCIe for LLMs: which models fit?

Cards needed for 50B+ models released since 2026, with vLLM tensor parallelism at 32K context

ModelSizeCardsTotal priceTotal VRAMSpeed t/sTotal power
Qwen3.8 Flash Next176B (6B active)111.3 GB2 cardsNo price160 GB146700 W
DeepSeek V4 Flash284B (13B active)155.1 GB4 cardsNo price320 GB112.51,400 W
Hy3295B (21B active)182.2 GB4 cardsNo price320 GB611,400 W
GLM 5.3 Flash320B (18B active)199.7 GB4 cardsNo price320 GB851,400 W
GLM 5.3744B (40B active)467.3 GB8 cardsNo price640 GB45.32,800 W
Hy4 preview770B (49B active)467.3 GB8 cardsNo price640 GB39.72,800 W
DeepSeek V4.1 Flash763B (8 / 16B active)510.3 GB8 cardsEngram in RAMNo price640 GB75.12,800 W
Kimi K2.7 Code1,000B (32B active)583.7 GB8 cardsNo price640 GB57.22,800 W

Card counts are 1 / 2 / 4 / 8, each using 90% of its VRAM. Speed is based on one card’s bandwidth; price covers GPUs only. Rows marked “Engram in RAM” need about 203 GB of extra system RAM. Supports NVLink (600 GB/s). Whether it runs also depends on vLLM support.

H100 PCIe rental price per hour

On-demand rental of the H100 PCIe starts at $1.99 per hour (RunPod), about $484/month at 8 hours a day.

OptionPrice4 h/day8 h/day24 h/dayView
RunPodRentUpdated 42 min. ago$1.99/hr$242/mo$484/mo$1,453/moRent
Vast.aiRentUpdated 42 min. ago$2.59/hr$315/mo$630/mo$1,890/moRent

Other markets

eBay · United States
N/A
No listings
View
Amazon.jp · Japan
N/A
No listings
View

Buy links may carry affiliate tags; they do not change the price you pay.

FAQ

How large a model can the H100 PCIe run?

A single H100 PCIe has 80GB of HBM2e VRAM. At 4-bit quantization and 8K context it fits a dense model of up to about 136B parameters, or about 73B at 8-bit.

Can the H100 PCIe run a 70B model?

Yes. At 4-bit with an 8K context a 70B dense model needs about 43GB, which fits in the H100 PCIe's 80GB.

What are the best LLMs to run on the H100 PCIe?

Among models released in 2026, the largest dense model that fits entirely in VRAM at 4-bit is Gemma 4 31B, at an estimated 72.7 tokens/s. Among MoE models, Qwen3.6 35B A3B is the fastest at about 188.3 tokens/s. The full list is in the table above.

How many tokens per second does the H100 PCIe get on LLMs?

Estimated from memory bandwidth at 4-bit and 8K context: 8B dense: about 226.5 tokens/s, 14B dense: about 142.8 tokens/s, 32B dense: about 68.8 tokens/s, and 70B dense: about 33.1 tokens/s. MoE models read only the active parameters for each token, so they run much faster. Real-world results are usually 70%–100% of the estimate.

What LLMs can dual H100 PCIe cards run?

Two H100 PCIe cards (160GB total) with vLLM tensor parallelism at 4-bit and 32K context can fit Qwen3.8 Flash Next. To run Kimi K2.7 Code at 4-bit with 32K context, you need 8 cards.

How much does a H100 PCIe cost?

No price data yet; it will appear here once the feed is live.

How much does it cost to rent a H100 PCIe per hour?

As of Sep 29, 2026: RunPod at $1.99 per hour and Vast.ai at $2.59 per hour. All are single-GPU on-demand rates, excluding storage and data transfer.

H100 PCIe vs A100 80GB PCIe: which is better for LLMs?

The A100 80GB PCIe has 80GB of VRAM and 1,935 GB/s of bandwidth; the H100 PCIe has 80GB and 2,000 GB/s. For 70B dense models at 4-bit, the H100 PCIe is estimated at about 33.1 tokens/s and the A100 80GB PCIe at about 32 tokens/s.