Popular Open-Weight Models for Local Deployment
| Version | Vendor | License | ||||||
|---|---|---|---|---|---|---|---|---|
| Alibaba | Apr 27, 2025 | 8.2 | 8.2 | 7 | GeForce RTX 4080— · 16 GB | 13,232,997 | Apache-2.0 | |
| Alibaba | Feb 27, 2026 | 9.7 | 9.7 | 8 | GeForce RTX 4080— · 16 GB | 12,375,273 | Apache-2.0 | |
| Google DeepMind | Mar 11, 2026 | 31.3 | 31.3 | 26 | Arc Pro B65— · 32 GB | 8,332,852 | Apache-2.0 | |
| Google DeepMind | Mar 11, 2026 | 25.8 | 4 | 6plus 13 GB of RAM | Radeon RX 7900 XT— · 20 GB | 8,211,474 | Apache-2.0 | |
| gpt-oss 20B | OpenAI | Aug 4, 2025 | 20.9 | 3.6 | 4plus 12 GB of RAM | GeForce RTX 4080— · 16 GB | 6,448,506 | Apache-2.0 |
| Alibaba | Aug 5, 2026 | 27.8 | 27.8 | 19 | Radeon RX 7900 XT— · 20 GB | 5,739,341 | Apache-2.0 | |
| Meta AI | Jul 18, 2024 | 8 | 8 | 7 | GeForce RTX 4080— · 16 GB | 5,734,979 | Llama 3.1 Community License | |
| Alibaba | Apr 21, 2026 | 27.8 | 27.8 | 19 | Radeon RX 7900 XT— · 20 GB | 5,366,737 | Apache-2.0 | |
| gpt-oss 120B | OpenAI | Aug 4, 2025 | 116.8 | 5.1 | 4plus 65 GB of RAM | RTX PRO 5000 Blackwell 72GB— · 72 GB | 5,259,501 | Apache-2.0 |
| Alibaba | Apr 27, 2025 | 32.8 | 32.8 | 22 | Arc Pro B60— · 24 GB | 5,024,271 | Apache-2.0 | |
| Google DeepMind | Mar 2, 2026 | 8 | 8 | 7 | GeForce RTX 4080— · 16 GB | 4,850,749 | Apache-2.0 | |
| DeepSeek | Jul 31, 2026 | 284 | 13 | 7plus 155 GB of RAM | Mac Studio M3 Ultra 256GB— · 256 GB | 4,559,659 | MIT | |
| Alibaba | Apr 15, 2026 | 36 | 3 | 4plus 19 GB of RAM | Arc Pro B60— · 24 GB | 4,546,612 | Apache-2.0 | |
| Google DeepMind | May 23, 2026 | 12 | 12 | 11 | GeForce RTX 4080— · 16 GB | 3,195,490 | Apache-2.0 | |
| Mistral AI | May 22, 2024 | 7.3 | 7.3 | 7 | GeForce RTX 4080— · 16 GB | 2,648,636 | Apache-2.0 | |
| Moonshot AI | Jun 13, 2026 | 2,800 | 104 | 33plus 1,531 GB of RAM | A100 40GB PCIe— · 40 GB · Experts offloaded | 2,639,566 | Kimi K3 License | |
| Z.ai / Zhipu AI | Jan 19, 2026 | 30 | 3 | 4plus 17 GB of RAM | Radeon RX 7900 XT— · 20 GB | 1,935,018 | MIT | |
| DeepSeek | Dec 1, 2025 | 671 | 37 | 12plus 365 GB of RAM | Mac Studio M3 Ultra 512GB— · 512 GB | 1,559,697 | MIT | |
| Meta AI | Sep 18, 2024 | 3.2 | 3.2 | 4 | GeForce RTX 4080— · 16 GB | 1,419,885 | Llama 3.2 Community License | |
| Z.ai / Zhipu AI | Jun 16, 2026 | 744 | 40 | 41plus 406 GB of RAM | Mac Studio M3 Ultra 512GB— · 512 GB | 1,134,389 | MIT | |
| DeepSeek | May 29, 2025 | 8.2 | 8.2 | 7 | GeForce RTX 4080— · 16 GB | 1,003,991 | MIT | |
| Meta AI | Nov 26, 2024 | 70.6 | 70.6 | 43 | GeForce RTX 4090 48GB (modded)— · 48 GB | 815,592 | Llama 3.3 Community License | |
| Z.ai / Zhipu AI | Aug 25, 2026 | 320 | 18 | 30plus 174 GB of RAM | Mac Studio M3 Ultra 256GB— · 256 GB | 654,957 | MIT | |
| Moonshot AI | Apr 14, 2026 | 1,000 | 32 | 9plus 552 GB of RAM | GeForce RTX 4080— · 16 GB · Experts offloaded | 638,370 | Modified MIT | |
| DeepSeek | Jan 20, 2025 | 32.8 | 32.8 | 22 | Arc Pro B60— · 24 GB | 562,818 | MIT | |
| Moonshot AI | Jan 1, 2026 | 1,000 | 32 | 9plus 552 GB of RAM | GeForce RTX 4080— · 16 GB · Experts offloaded | 518,180 | Modified MIT | |
| DeepSeek | Jan 20, 2025 | 14.8 | 14.8 | 11 | GeForce RTX 4080— · 16 GB | 396,404 | MIT | |
| Mistral AI | Jul 17, 2024 | 12.3 | 12.3 | 10 | GeForce RTX 4080— · 16 GB | 358,821 | Apache-2.0 | |
| Z.ai / Zhipu AI | Aug 25, 2026 | 744 | 40 | 41plus 406 GB of RAM | Mac Studio M3 Ultra 512GB— · 512 GB | 303,534 | GLM-5.3 License | |
| Mistral AI | Oct 31, 2025 | 14 | 14 | 11 | GeForce RTX 4080— · 16 GB | 266,145 | Apache-2.0 | |
| Moonshot AI | Oct 30, 2025 | 49.1 | 3 | 3plus 27 GB of RAM | Arc Pro B65— · 32 GB | 187,057 | MIT | |
| Meta AI | Apr 2, 2025 | 108.6 | 17 | 10plus 55 GB of RAM | Instinct MI210— · 64 GB | 174,923 | Llama 4 Community License | |
| Mistral AI | Oct 31, 2025 | 8.9 | 8.9 | 8 | GeForce RTX 4080— · 16 GB | 142,610 | Apache-2.0 | |
| Mistral AI | Jun 19, 2025 | 24 | 24 | 16 | GeForce RTX 4080— · 16 GB | 126,611 | Apache-2.0 | |
| Z.ai / Zhipu AI | Jul 20, 2025 | 106 | 12 | 7plus 56 GB of RAM | Instinct MI210— · 64 GB | 114,264 | MIT | |
| gpt-oss Safeguard 20B | OpenAI | Sep 18, 2025 | 20.9 | 3.6 | 4plus 12 GB of RAM | GeForce RTX 4080— · 16 GB | 78,266 | Apache-2.0 |
| Moonshot AI | Sep 3, 2025 | 1,000 | 32 | 9plus 552 GB of RAM | GeForce RTX 4080— · 16 GB · Experts offloaded | 37,308 | Modified MIT |
Minimum VRAM = Q4_K_M weights + 8K context + runtime overhead · Used prices are median asking prices, not sold prices · Methodology