GPUs tracked56Price history27 daysPrice points (24h)30,049 pointsBiggest 24h dropGeForce RTX 4060 Ti 16GB-16.4%Biggest 24h riseRadeon RX 7800 XT+19.1%Updated
$ USD
English

RTX A6000 Price & Used Prices: Which Local LLMs Can It Run?

The RTX A6000 has 48GB of GDDR6 ECC memory with 768 GB/s of bandwidth. The median asking price is about $8,877 on Amazon.jp. At 4-bit quantization and 8K context it fits a dense model of up to about 79B parameters, at roughly 12 tokens/s.

RTX A6000 price history

Price updated
$14,032$11,723$9,413$7,103$4,793
Low $6,113High $12,713
Sep 22Sep 26Sep 29

Buy links may carry affiliate tags; they do not change the price you pay.

Full RTX A6000 specifications

  • Memory48 GB · GDDR6 ECC
  • Bandwidth768 GB/s
  • FP16 / BF16154.9 TFLOPS
  • INT8309.7 TOPS
Bus width
384-bit
Architecture
Ampere (GA102)
Cores
10,752
Tensor cores
336
Ecosystem
CUDA
Interface
PCIe 4.0 x16
NVLink
112.5 GB/s
Power
300 W
Power connector
8-pin EPS
Slots
2
Length
267 mm
Type
Workstation
Launch date
Oct 5, 2020

What LLMs can the RTX A6000 run?

Models up to 50B parameters on a single RTX A6000, versions released since 2026 only

ModelSizeRuns?Speed t/sMax context
Gemma 4 E4B8B5 GBRuns102.1128K
Qwen3.5 9B9.7B5.7 GBRuns89.8256K
Gemma 4 12B12B7.1 GBRuns72.5256K
Gemma 4 26B A4B25.2B (3.8B active)16.9 GBRuns113.1256K
Qwen3.6 27B27.8B16.8 GBRuns32.2256K
Qwen3.8 27B27.8B16.5 GBRuns32.8256K
GLM 4.7 Flash30B (3B active)18.3 GBRuns131.1198K
Gemma 4 31B30.7B18.3 GBRuns29256K
Qwen3.6 35B A3B36B (3B active)22.1 GBRuns134.9256K

RunsFits fully in VRAM; speed is estimated from VRAM bandwidth.

Needs RAM offloadMoE only: expert weights sit in system RAM, so speed is estimated from 70 GB/s RAM bandwidth.

Too bigWon’t fit on one card at this quantization, even with RAM offload.

Dual, 4x and 8x RTX A6000 for LLMs: which models fit?

Cards needed for 50B+ models released since 2026, with vLLM tensor parallelism at 32K context

ModelSizeCardsTotal priceTotal VRAMSpeed t/sTotal power
Qwen3.8 Flash Next176B (6B active)111.3 GB4 cards$35,507192 GB87.61,200 W
DeepSeek V4 Flash284B (13B active)155.1 GB4 cards$35,507192 GB59.81,200 W
Hy3295B (21B active)182.2 GB8 cards$71,013384 GB27.62,400 W
GLM 5.3 Flash320B (18B active)199.7 GB8 cards$71,013384 GB41.32,400 W

Card counts are 1 / 2 / 4 / 8, each using 90% of its VRAM. Speed is based on one card’s bandwidth; price covers GPUs only. Supports NVLink (112.5 GB/s). Whether it runs also depends on vLLM support.

RTX A6000 rental price per hour

On-demand rental of the RTX A6000 starts at $0.33 per hour (RunPod), about $80/month at 8 hours a day. Buying one (Amazon.jp median $8,877) pays off in about 22 years, 11 years or 3.7 years at 4, 8 or 24 hours a day.

OptionPrice4 h/day8 h/day24 h/dayView
RunPodRentUpdated 24 min. ago$0.33/hr$40/mo$80/mo$241/moRent
Vast.aiRentUpdated 24 min. ago$0.43/hr$52/mo$104/mo$313/moRent
Amazon.jpUpdated 6 hr. ago$8,877Payback ~22 yrsPower $7/moPayback ~11 yrsPower $13/moPayback ~3.7 yrsPower $39/moBuy

Break-even uses the US residential average ($0.18/kWh) and 300 W rated power, excluding resale value and the rest of the build.

Other markets

eBay · United States
$5,195
19 listings
View
Xianyu · Mainland China
$4,612
16 listings
View

Buy links may carry affiliate tags; they do not change the price you pay.

FAQ

How large a model can the RTX A6000 run?

A single RTX A6000 has 48GB of GDDR6 ECC VRAM. At 4-bit quantization and 8K context it fits a dense model of up to about 79B parameters, or about 43B at 8-bit.

Can the RTX A6000 run a 70B model?

Yes. At 4-bit with an 8K context a 70B dense model needs about 43GB, which fits in the RTX A6000's 48GB.

What are the best LLMs to run on the RTX A6000?

Among models released in 2026, the largest dense model that fits entirely in VRAM at 4-bit is Gemma 4 31B, at an estimated 29 tokens/s. Among MoE models, Qwen3.6 35B A3B is the fastest at about 134.9 tokens/s. The full list is in the table above.

How many tokens per second does the RTX A6000 get on LLMs?

Estimated from memory bandwidth at 4-bit and 8K context: 8B dense: about 97.9 tokens/s, 14B dense: about 59 tokens/s, 32B dense: about 27.3 tokens/s, and 70B dense: about 12.9 tokens/s. MoE models read only the active parameters for each token, so they run much faster. Real-world results are usually 70%–100% of the estimate.

What LLMs can dual RTX A6000 cards run?

With vLLM tensor parallelism at 4-bit and 32K context, two RTX A6000 cards cannot fit any model of 50B or larger released in 2026. At 4-bit with 32K context the minimum is 4 cards, which fits Qwen3.8 Flash Next and DeepSeek V4 Flash.

How much does a RTX A6000 cost?

As of Sep 29, 2026: eBay median $5,195 (19 listings), Amazon.jp median $8,877 (5 listings), and Xianyu median $4,612 (16 listings). Figures are medians of live listings, refreshed every 6 hours.

Is the RTX A6000 price going up or down?

As of Sep 29, 2026, the Amazon.jp median is down 1% over 7 days, down 6.7% over 30 days.

How much does it cost to rent a RTX A6000 per hour?

As of Sep 29, 2026: RunPod at $0.33 per hour and Vast.ai at $0.43 per hour. All are single-GPU on-demand rates, excluding storage and data transfer.

Is it cheaper to buy or rent a RTX A6000 at 8 hours a day?

Taking the Amazon.jp median of $8,877 and electricity at $0.18 per kWh, at 8 hours a day buying pays for itself in about 11 years (versus RunPod at $0.33 per hour). If you will use it for longer than 11 years, buy; otherwise, rent.

RTX A6000 vs RTX 6000 Ada: which is better for LLMs?

The RTX 6000 Ada has 48GB of VRAM and 960 GB/s of bandwidth; the RTX A6000 has 48GB and 768 GB/s. For 70B dense models at 4-bit, the RTX A6000 is estimated at about 12.9 tokens/s and the RTX 6000 Ada at about 16.1 tokens/s.