How Much VRAM Does Kimi K2 Need to Run Locally?
See devices that run itKimi K2 is a MoE model with 1,000B total parameters and 32B active per token. At Q4 it needs 559GB of VRAM to sit entirely on the GPU; with experts offloaded to system RAM it needs only 9GB of VRAM plus 552GB of RAM. No single card in the database fits it; it needs multiple GPUs or the unified memory of a Mac Studio. In offload mode the cheapest card that runs it is the GeForce RTX 2080 Ti, at roughly 4 tokens/s.
Kimi K2 details
VRAM needed for Kimi K2 by quantization and context
Total at 8KMoE: all 1,000B parameters must stay resident in memory, while speed is set by the 32B active per token.
| Quantization | WeightsGB | Total at 8KGB | Total at 32KGB | Total at 128KGB |
|---|---|---|---|---|
| Q4_K_M | 557.4est. | 558.9 | 560.5 | 567 |
| Q5_K_M | 674.7est. | 676.3 | 677.9 | 684.3 |
| Q8_0 | 1,036.6est. | 1,038.1 | 1,039.7 | 1,046.1 |
| FP16 | 1,955.8est. | 1,957.3 | 1,958.9 | 1,965.4 |
Total = weights + KV cache (FP16) + 1GB runtime overhead. Weights marked GGUF are measured file sizes from the repository; the rest are estimated from bytes per parameter. Contexts beyond this model's 256K limit show a dash.
Which GPUs can run Kimi K2?
Q4_K_M · 8K context · sorted by eBay used priceNo single card in the database fits Kimi K2: it needs at least 559GB of VRAM at Q4_K_M with an 8K context.
A multi-GPU build needs at least 559GB of VRAM in total.
None of the unified-memory machines in the database fit it either.
GPUs that run Kimi K2 with experts offloaded to RAM
Expert weights live in system RAM (needs at least 552 GB); VRAM holds only attention layers, shared experts and the KV cache, at least 9 GB at Q4. Speed is bound by RAM bandwidth (estimated at 70 GB/s) and is far slower than a full-VRAM setup.
Run Kimi K2 with Ollama, llama.cpp or vLLM
The Ollama library has no official tag for it yet.
vLLM serves the original-precision weights, which needs far more memory than GGUF and usually more than one GPU.
vllm serve moonshotai/Kimi-K2-Instruct-0905
File names follow that week's model-library snapshot; the definitions are in the methodology. Methodology
FAQ
How much VRAM does Kimi K2 need?
At Q4_K_M with an 8K context Kimi K2 needs about 559GB of VRAM; 672GB leaves comfortable headroom.
What is the cheapest GPU that runs Kimi K2?
No single card in the database fits Kimi K2; it needs multiple GPUs or a Mac Studio's unified memory.
Can you run Kimi K2 with Ollama?
Not yet: there is no official Ollama tag and no public GGUF build for it.