收录显卡45价格追踪8 天自 2026年9月3日近 24h 价格数据7,968 条24h 降价最多Mac Studio M4 Max 128GB-6.7%24h 涨价最多GeForce RTX 5070 Ti+21.0%更新于

本地部署 Llama 需要什么显卡?各版本显存需求一览

Llama 目前有 4 个能本地部署的版本,最小的 Llama 3.2 3B 在 Q4 量化下需要 4GB 显存,最大的 Llama 4 Scout 17B 16E 需要 10GB。最便宜能跑起来的设备是 Tesla V100 16GB,约 ¥880

Llama 各版本的显存需求

闲鱼
共 4 个版本
版本发布日期上下文
Llama 3.2 3B稠密3.2B2024年9月18日128K4Tesla V100 16GB¥880 · 16 GB
Llama 3.1 8B稠密8B2024年7月18日128K7Tesla V100 16GB¥880 · 16 GB
Llama 3.3 70B稠密70.6B2024年11月26日128K43CMP 170HX 64GB¥14,050 · 64 GB
Llama 4 Scout 17B 16EMoE108.6B(激活 17B)2025年4月2日10M10另需 55 GB 内存Tesla V100 16GB¥880 · 16 GB · 专家卸载

最低显存 = Q4_K_M 权重 + 8K 上下文 KV cache + 1GB 运行开销,向上取整;最便宜能跑的设备按当前价格基准的中位价升序,取第一台装得下的。

关于 Llama

Llama 由 Meta AI 发布,许可证为 Llama 3.2 Community License / Llama 3.1 Community License / Llama 3.3 Community License / Llama 4 Community License。本页收录 4 个版本,参数量从 3.2B 到 108.6B,结构参数与量化体积依据 Hugging Face 等来源核验后更新。

厂商Meta AI
许可证Llama 3.2 Community License / Llama 3.1 Community License / Llama 3.3 Community License / Llama 4 Community License
Hugging Face 组织meta-llama