GPUs tracked56Price history27 daysPrice points (24h)34,228 pointsBiggest 24h dropA100 80GB PCIe-20.0%Biggest 24h riseRTX PRO 5000 48GB+17.5%Updated
$ USD
English

GeForce RTX 4060 Ti 16GB Price & Used Prices: Which Local LLMs Can It Run?

The GeForce RTX 4060 Ti 16GB has 16GB of GDDR6 memory with 288 GB/s of bandwidth. The median asking price is about $537 on Xianyu. At 4-bit quantization and 8K context it fits a dense model of up to about 23B parameters, at roughly 14 tokens/s.

GeForce RTX 4060 Ti 16GB price history

Price updated
$560$548$536$524$512
Low $521High $550
Sep 24Sep 26Sep 29

Buy links may carry affiliate tags; they do not change the price you pay.

Full GeForce RTX 4060 Ti 16GB specifications

  • Memory16 GB · GDDR6
  • Bandwidth288 GB/s
  • FP16 / BF1688.3 TFLOPS
  • FP8 / INT8176.5 / 176.5 TOPS
Bus width
128-bit
Architecture
Ada Lovelace (AD106)
Cores
4,352
Tensor cores
136
Ecosystem
CUDA
Interface
PCIe 4.0 x8
NVLink
Not supported
Power
165 W
Power connector
1× 8-pin / 16-pin 12VHPWR
Type
Consumer
Launch date
Jul 18, 2023
MSRP
$499

What LLMs can the GeForce RTX 4060 Ti 16GB run?

Models up to 50B parameters on a single GeForce RTX 4060 Ti 16GB, versions released since 2026 only

ModelSizeRuns?Speed t/sMax context
Gemma 4 E4B8B5 GBRuns40.4128K
Qwen3.5 9B9.7B5.7 GBRuns35.3256K
Gemma 4 12B12B7.1 GBRuns28.2256K
Gemma 4 26B A4B25.2B (3.8B active)16.9 GBNeeds RAM offload32.3256K
Qwen3.6 27B27.8B16.8 GBToo big——
Qwen3.8 27B27.8B16.5 GBToo big——
GLM 4.7 Flash30B (3B active)18.3 GBNeeds RAM offload33.5198K
Gemma 4 31B30.7B18.3 GBToo big——
Qwen3.6 35B A3B36B (3B active)22.1 GBNeeds RAM offload44256K

RunsFits fully in VRAM; speed is estimated from VRAM bandwidth.

Needs RAM offloadMoE only: expert weights sit in system RAM, so speed is estimated from 70 GB/s RAM bandwidth.

Too bigWon’t fit on one card at this quantization, even with RAM offload.

Dual, 4x and 8x GeForce RTX 4060 Ti 16GB for LLMs: which models fit?

Cards needed for 50B+ models released since 2026, with vLLM tensor parallelism at 32K context

ModelSizeCardsTotal priceTotal VRAMSpeed t/sTotal power
Qwen3.8 Flash Next176B (6B active)111.3 GB8 cards$4,297128 GB421,320 W

Card counts are 1 / 2 / 4 / 8, each using 90% of its VRAM. Speed is based on one card’s bandwidth; price covers GPUs only. No NVLink; cards talk over PCIe. Whether it runs also depends on vLLM support.

Other markets

eBay · United States
$700
5 listings
View
Amazon.jp · Japan
$1,327
3 listings
View

Buy links may carry affiliate tags; they do not change the price you pay.

FAQ

How large a model can the GeForce RTX 4060 Ti 16GB run?

A single GeForce RTX 4060 Ti 16GB has 16GB of GDDR6 VRAM. At 4-bit quantization and 8K context it fits a dense model of up to about 23B parameters, or about 13B at 8-bit. MoE models can offload experts to system RAM to run larger ones, such as Qwen3.6 35B A3B (36B).

Can the GeForce RTX 4060 Ti 16GB run a 70B model?

Not on a single card. At 4-bit with an 8K context a 70B dense model needs about 43GB, more than the GeForce RTX 4060 Ti 16GB's 16GB. It runs if you split the layers across 3 cards with llama.cpp (48GB total).

What are the best LLMs to run on the GeForce RTX 4060 Ti 16GB?

Among models released in 2026, the largest dense model that fits entirely in VRAM at 4-bit is Gemma 4 12B, at an estimated 28.2 tokens/s. The full list is in the table above.

How many tokens per second does the GeForce RTX 4060 Ti 16GB get on LLMs?

Estimated from memory bandwidth at 4-bit and 8K context: 8B dense: about 38.6 tokens/s and 14B dense: about 22.8 tokens/s. MoE models read only the active parameters for each token, so they run much faster. Real-world results are usually 70%–100% of the estimate.

What LLMs can dual GeForce RTX 4060 Ti 16GB cards run?

With vLLM tensor parallelism at 4-bit and 32K context, two GeForce RTX 4060 Ti 16GB cards cannot fit any model of 50B or larger released in 2026. At 4-bit with 32K context the minimum is 8 cards, which fits Qwen3.8 Flash Next.

How much does a GeForce RTX 4060 Ti 16GB cost?

As of Sep 29, 2026: eBay median $700 (5 listings), Amazon.jp median $1,327 (3 listings), and Xianyu median $537 (3 listings). Figures are medians of live listings, refreshed every 6 hours.

Is the GeForce RTX 4060 Ti 16GB price going up or down?

As of Sep 29, 2026, the Xianyu median is down 2.1% over 7 days, down 2.1% over 30 days.

GeForce RTX 4060 Ti 16GB vs Tesla T10: which is better for LLMs?

The Tesla T10 has 16GB of VRAM and 403 GB/s of bandwidth; the GeForce RTX 4060 Ti 16GB has 16GB and 288 GB/s. At the Xianyu median, the Tesla T10 costs about $199, 63% less than the GeForce RTX 4060 Ti 16GB ($537). For 14B dense models at 4-bit, the GeForce RTX 4060 Ti 16GB is estimated at about 22.8 tokens/s and the Tesla T10 at about 31.7 tokens/s.