What Do You Need to Run gpt-oss Locally? VRAM by Version
gpt-oss currently has 3 versions you can deploy locally. The smallest, gpt-oss 20B, needs 4GB of VRAM at Q4; the largest, gpt-oss 120B, needs 4GB. The cheapest device that gets one running is the GeForce RTX 4080, with no used-price data yet.
VRAM needed for each gpt-oss version
3 versions| Version | Released | Context | ||||
|---|---|---|---|---|---|---|
| gpt-oss 20BMoE | 20.9B (3.6B active) | Aug 4, 2025 | 128K | 4plus 12 GB of RAM | GeForce RTX 4080— · 16 GB | 6,448,506 |
| gpt-oss Safeguard 20BMoE | 20.9B (3.6B active) | Sep 18, 2025 | 128K | 4plus 12 GB of RAM | GeForce RTX 4080— · 16 GB | 78,266 |
| gpt-oss 120BMoE | 116.8B (5.1B active) | Aug 4, 2025 | 128K | 4plus 65 GB of RAM | RTX PRO 5000 Blackwell 72GB— · 72 GB | 5,259,501 |
Minimum VRAM = Q4_K_M weights + KV cache for an 8K context + 1GB runtime overhead, rounded up. The cheapest device is the first one that fits when sorting by median used price on the current price basis.
About gpt-oss
gpt-oss is published by OpenAI under Apache-2.0. This page covers 3 versions ranging from 20.9B to 116.8B parameters; structure parameters and quantized sizes are synced weekly from Hugging Face.
VendorOpenAI
LicenseApache-2.0
Official sitehttps://openai.com/open-models/
Hugging Face orgopenai