What Do You Need to Run gpt-oss Locally? VRAM by Version
gpt-oss currently has 3 versions you can deploy locally. The smallest, gpt-oss 20B, needs 4GB of VRAM at Q4; the largest, gpt-oss 120B, needs 4GB. The cheapest device that gets one running is the Tesla V100 16GB, at about $131.
VRAM needed for each gpt-oss version
Xianyu
| Version | Released | Context | |||
|---|---|---|---|---|---|
| gpt-oss 20BMoE | 20.9B (3.6B active) | Aug 4, 2025 | 128K | 4plus 12 GB of RAM | Tesla V100 16GB$131 · 16 GB |
| gpt-oss Safeguard 20BMoE | 20.9B (3.6B active) | Sep 18, 2025 | 128K | 4plus 12 GB of RAM | Tesla V100 16GB$131 · 16 GB |
| gpt-oss 120BMoE | 116.8B (5.1B active) | Aug 4, 2025 | 128K | 4plus 65 GB of RAM | Tesla V100 16GB$131 · 16 GB · Experts offloaded |
Minimum VRAM = Q4_K_M weights + KV cache for an 8K context + 1GB runtime overhead, rounded up. The cheapest device is the first one that fits when sorting by median price on the current price basis.
About gpt-oss
gpt-oss is published by OpenAI under Apache-2.0. This page covers 3 versions ranging from 20.9B to 116.8B parameters. New official releases are scanned weekly; architecture parameters and sourced or estimated weight sizes are reviewed before publication.
VendorOpenAI
LicenseApache-2.0
Official sitehttps://openai.com/open-models/
Hugging Face orgopenai