收录显卡41价格追踪3 天自 2026年9月3日近 24h 价格数据11,753 条24h 降价最多Radeon RX 7900 XTX-5.6%24h 涨价最多A100 40GB PCIe+48.1%数据更新于 2026年9月5日

本地部署 Llama 需要什么显卡?各版本显存需求一览

Llama 目前有 4 个能本地部署的版本,最小的 Llama 3.2 3B 在 Q4 量化下需要 4GB 显存,最大的 Llama 4 Scout 17B 16E 需要 10GB。最便宜能跑起来的设备是 GeForce RTX 2080 Ti,二手约 ¥1,680

Llama 各版本的显存需求

共 4 个版本
版本发布日期上下文
Llama 3.2 3B稠密3.2B2024年9月18日128K4GeForce RTX 2080 Ti¥1,680 · 11 GB1,419,885
Llama 3.1 8B稠密8B2024年7月18日128K7GeForce RTX 2080 Ti¥1,680 · 11 GB5,734,979
Llama 3.3 70B稠密70.6B2024年11月26日128K43CMP 170HX 64GB (改装)¥12,850 · 64 GB815,592
Llama 4 Scout 17B 16EMoE108.6B(激活 17B)2025年4月2日10M10另需 55 GB 内存GeForce RTX 2080 Ti¥1,680 · 11 GB · 专家卸载174,923

最低显存 = Q4_K_M 权重 + 8K 上下文 KV cache + 1GB 运行开销,向上取整;最便宜能跑的设备按当前价格基准的二手中位价升序,取第一台装得下的。

关于 Llama

Llama 由 Meta AI 发布,许可证为 Llama 3.2 Community License / Llama 3.1 Community License / Llama 3.3 Community License / Llama 4 Community License。本页收录 4 个版本,参数量从 3.2B 到 108.6B,结构参数与量化体积每周从 Hugging Face 同步。

厂商Meta AI
许可证Llama 3.2 Community License / Llama 3.1 Community License / Llama 3.3 Community License / Llama 4 Community License
Hugging Face 组织meta-llama