本地部署 Llama 需要什么显卡?各版本显存需求一览
Llama 目前有 4 个能本地部署的版本,最小的 Llama 3.2 3B 在 Q4 量化下需要 4GB 显存,最大的 Llama 4 Scout 17B 16E 需要 10GB。最便宜能跑起来的设备是 GeForce RTX 2080 Ti,二手约 ¥1,680。
Llama 各版本的显存需求
共 4 个版本| 版本 | 发布日期 | 上下文 | ||||
|---|---|---|---|---|---|---|
| Llama 3.2 3B稠密 | 3.2B | 2024年9月18日 | 128K | 4 | GeForce RTX 2080 Ti¥1,680 · 11 GB | 1,419,885 |
| Llama 3.1 8B稠密 | 8B | 2024年7月18日 | 128K | 7 | GeForce RTX 2080 Ti¥1,680 · 11 GB | 5,734,979 |
| Llama 3.3 70B稠密 | 70.6B | 2024年11月26日 | 128K | 43 | CMP 170HX 64GB (改装)¥12,850 · 64 GB | 815,592 |
| Llama 4 Scout 17B 16EMoE | 108.6B(激活 17B) | 2025年4月2日 | 10M | 10另需 55 GB 内存 | GeForce RTX 2080 Ti¥1,680 · 11 GB · 专家卸载 | 174,923 |
最低显存 = Q4_K_M 权重 + 8K 上下文 KV cache + 1GB 运行开销,向上取整;最便宜能跑的设备按当前价格基准的二手中位价升序,取第一台装得下的。
关于 Llama
Llama 由 Meta AI 发布,许可证为 Llama 3.2 Community License / Llama 3.1 Community License / Llama 3.3 Community License / Llama 4 Community License。本页收录 4 个版本,参数量从 3.2B 到 108.6B,结构参数与量化体积每周从 Hugging Face 同步。
厂商Meta AI
许可证Llama 3.2 Community License / Llama 3.1 Community License / Llama 3.3 Community License / Llama 4 Community License
Hugging Face 组织meta-llama