How Much VRAM Does Kimi K3 Need to Run Locally?
See devices that run itKimi K3 is a MoE model with 2,800B total parameters and 104B active per token. At Q4 it needs 1,562GB of VRAM to sit entirely on the GPU; with experts offloaded to system RAM it needs only 33GB of VRAM plus 1,531GB of RAM. No single card in the database fits it; it needs multiple GPUs or the unified memory of a Mac Studio. In offload mode the cheapest card that runs it is the A100 40GB PCIe, at roughly 2 tokens/s.
Kimi K3 details
VRAM needed for Kimi K3 by quantization and context
Total at 8KMoE: all 2,800B parameters must stay resident in memory, while speed is set by the 104B active per token.
| Quantization | WeightsGB | Total at 8KGB | Total at 32KGB | Total at 128KGB |
|---|---|---|---|---|
| Q4_K_M | 1,560.7est. | 1,561.9 | 1,562.6 | 1,565.1 |
| Q5_K_M | 1,889.3est. | 1,890.5 | 1,891.1 | 1,893.7 |
| Q8_0 | 2,902.4est. | 2,903.6 | 2,904.2 | 2,906.7 |
| FP16 | 5,476.2est. | 5,477.4 | 5,478 | 5,480.6 |
Total = weights + KV cache (FP16) + 1GB runtime overhead. Weights marked GGUF are measured file sizes from the repository; the rest are estimated from bytes per parameter. Contexts beyond this model's 1M limit show a dash.
Which GPUs can run Kimi K3?
Q4_K_M · 8K context · sorted by eBay used priceNo single card in the database fits Kimi K3: it needs at least 1,562GB of VRAM at Q4_K_M with an 8K context.
A multi-GPU build needs at least 1,562GB of VRAM in total.
None of the unified-memory machines in the database fit it either.
GPUs that run Kimi K3 with experts offloaded to RAM
Expert weights live in system RAM (needs at least 1,531 GB); VRAM holds only attention layers, shared experts and the KV cache, at least 33 GB at Q4. Speed is bound by RAM bandwidth (estimated at 70 GB/s) and is far slower than a full-VRAM setup.
| Device | VRAMGB | Used price | RAM neededGB | Est. t/s |
|---|---|---|---|---|
| 40 | — | 1,531 | 1.7 | |
| 40 | — | 1,531 | 1.7 | |
| 48 | — | 1,531 | 1.6 | |
| 48 | — | 1,531 | 1.5 | |
| 48 | — | 1,531 | 1.6 | |
| 64 | — | 1,531 | 1.7 | |
| 64 | — | 1,531 | 1.6 | |
| 72 | — | 1,531 | 1.6 | |
| 84 | — | 1,531 | 1.7 | |
| 96 | — | 1,531 | 1.7 | |
| Mac Studio M4 Max 128GB | 128 | — | 1,531 | 1.5 |
| Mac Studio M5 Max 128GB | 128 | — | 1,531 | 1.5 |
| MacBook Pro M4 Max 128GB | 128 | — | 1,531 | 1.5 |
| MacBook Pro M5 Max 128GB | 128 | — | 1,531 | 1.5 |
| Mac Studio M3 Ultra 512GB | 512 | — | 1,531 | 1.6 |
Run Kimi K3 with Ollama, llama.cpp or vLLM
The Ollama library has no official tag for it yet.
vLLM serves the original-precision weights, which needs far more memory than GGUF and usually more than one GPU.
vllm serve moonshotai/Kimi-K3
File names follow that week's model-library snapshot; the definitions are in the methodology. Methodology
FAQ
How much VRAM does Kimi K3 need?
At Q4_K_M with an 8K context Kimi K3 needs about 1,562GB of VRAM; 1,876GB leaves comfortable headroom.
What is the cheapest GPU that runs Kimi K3?
No single card in the database fits Kimi K3; it needs multiple GPUs or a Mac Studio's unified memory.
Can you run Kimi K3 with Ollama?
Not yet: there is no official Ollama tag and no public GGUF build for it.