Best Budget GPU for Local LLM Inference in 2026: The Cheapest Card at Every VRAM Tier

Translated from the 简体中文 edition · Read the original

On a tight budget, card shopping needs one number: price per GB of VRAM, the used price divided by capacity. It compares across generations, brands, consumer and workstation cards alike. As of September 5, 2026 the cheapest card per tier is: RX 7800 XT at 16GB, RX 7900 XTX at 24GB, Tesla V100 32GB at 32GB (not recommended, see below), Radeon PRO W7900 at 48GB, the CMP 170HX mod and Instinct MI210 at 64GB, and the Mac Studio M4 Max at 128GB. The table updates live; the rest of this guide covers which cheap cards to skip and how to buy used.

Illustration: a row of used graphics cards from different generations on a wooden table with blank price tags
Illustration · a row of used graphics cards from different generations on a wooden table with blank price tags

Budget cards come down to one number: price per GB

Price per GB defines "cheap" precisely: how many gigabytes of VRAM a dollar buys, and therefore which model tier you can run. Two conditions apply:

  1. Bandwidth cannot be too low: VRAM decides whether a model runs, bandwidth decides how fast. Below 500 GB/s, models above 14B feel sluggish, so 256 to 273 GB/s machines like the DGX Spark and Strix Halo are excluded from this ranking.
  2. The ecosystem must work: CUDA is painless, ROCm works on RDNA 3 / 4 (see the state of AMD), oneAPI (Intel Arc) runs llama.cpp with thin documentation, and old architectures (Volta, Pascal) are losing software support.

The home page ranking is already sorted by price per GB; this guide slices it by tier.

The cheapest card at every VRAM tier

The cheapest and second-cheapest card in six tiers, ranked live by current used median:

VRAM tierCheapest cardAmazonJPPrice per GBBandwidth GB/sNext cheapestAmazonJP
16 GBTesla V100 16GB$659$41.2900Radeon RX 9070 XT$884
24 GBArc Pro B60$1,164$48.5456Radeon RX 7900 XTX$1,738
32 GBTesla V100 32GB$1,141$35.7900Arc Pro B70$2,074
48 GBNo priced card in this tier
64 GBNo priced card in this tier
128 GBRyzen AI Max+ 395 128GB$3,913$30.6256DGX Spark 128GB$8,223

A tier covers cards from that capacity up to the next tier. Only cards with a current price are ranked.

Tier by tier:

  • 16GB: the RX 7800 XT has the lowest price per GB; its 624 GB/s runs an 8B model near 57 tokens/s, which is fine. For CUDA look at a used RTX 4080. The catalog now covers the 22GB modified RTX 2080 Ti in the next capacity tier.
  • 24GB: the RX 7900 XTX versus RTX 3090 question, see the state of AMD and the 3090 guide. This is the budget sweet spot: it runs 32B at 4-bit.
  • 32GB: the Tesla V100 32GB is astonishingly cheap per GB, and the next section explains why to skip it; the real choices are the Radeon AI PRO R9700, the Arc Pro B70, or stretching to an RTX 5090.
  • 48GB: the Radeon PRO W7900 has a warranty, two slots and 864 GB/s, steadier than a modded 48GB 4090; two 3090s are the other route, see multi-GPU basics.
  • 64GB: the CMP 170HX 64GB mod (HBM2e, 1.5 TB/s, PCIe x4) and the Instinct MI210 are enthusiast options, passively cooled with no display outputs.
  • 128GB: the Mac Studio M4 Max is the only 128GB device that works out of the box, see Mac unified memory.

Cheap cards to avoid

Three kinds of card are not worth it at any price per GB: old-architecture data-center cards, passively cooled accelerators, and modded cards without warranty.

ModelVRAM GBBandwidth GB/sINT8 TOPSPower WAmazonJPPrice per GB
Tesla V100 32GB32900250$1,141$35.7
Tesla V100 16GB16900250$659$41.2
GeForce RTX 2080 Ti 22GB22616215250
Instinct MI100321,22992300
CMP 170HX 40GB401,560404250
CMP 170HX 64GB641,493404250
GeForce RTX 4090 48GB481,008661450
GeForce RTX 4080 SUPER 32GB32736418320
  • Tesla V100 16GB / 32GB: Volta has FP16 tensor cores only, no BF16 or INT8, so many of llama.cpp's newer kernels take the slow path; PCIe 3.0; passive cooling that needs your own ducting; CUDA 13 dropped Volta, pinning you to old drivers. 900 GB/s of HBM2 looks great on paper for a 2017 card.
  • RTX 2080 Ti 22GB (modified): upgraded from the original 11GB. Turing lacks native BF16. Memory capacity increases, but bandwidth and architecture remain unchanged; check memory stability and the seller’s support for the modification.
  • Instinct MI100: first-generation CDNA, ROCm support in maintenance, half-rate BF16.
  • CMP 170HX mods: a mining die with reflowed HBM2e; PCIe 1.1 x4 makes model loading take minutes, no display output, no warranty, nobody to repair it.
  • RTX 4090 48GB / 4080 SUPER 32GB mods: real performance and real memory; the risk is the lifetime of the reflowed chips and where the BIOS came from, with no warranty at all. If you accept that, they are the cheapest single 48GB card.
Illustration: an open shipping box on a desk with a used graphics card in an anti-static bag
Illustration: an open shipping box on a desk with a used graphics card in an anti-static bag

How to buy used without getting burned

On eBay, lean on Top Rated sellers and 30-day returns; on Xianyu, on the platform inspection service and recorded tests; on both, retest yourself.

eBay:

  1. Choose Top Rated Plus sellers with returns accepted for 30 days; for private sellers, check the feedback history for previous GPU sales.
  2. Skip any listing whose title or description says for parts, as-is, artifacts or no display.
  3. On arrival check nvidia-smi -q for a power limit and BIOS version that match the model, then run memtest_vulkan.
  4. Import duties and shipping are not in this site's prices; add them for your country.

Xianyu (for readers buying in China):

  1. Only sellers who offer the platform inspection service, which checks the card before releasing payment.
  2. Ask for a 30-minute memtest_vulkan or OCCT VRAM test on video with the day's date visible.
  3. Treat "mining bundle", "no returns" and other low-priced listings as mining cards and expect another 20% off before considering them.
  4. Run your own 30-minute full-load test within 24 hours of delivery and file a return at the first artifact.

Verdict: what to buy at each tier

One answer per tier: RX 7800 XT or a used RTX 4080 at 16GB, RX 7900 XTX or RTX 3090 at 24GB, Radeon PRO W7900 or two 3090s at 48GB, Mac Studio M4 Max at 128GB; the 32GB and 64GB tiers have no painless cheap card.

  • The first budget priority is reaching 24GB: it is the threshold for 32B models and the sweet spot for price per GB.
  • Do not buy old architectures or modded cards for the unit price unless you can repair them yourself.
  • Run the three-year math before buying, see what running LLMs locally really costs.

FAQ

What is the cheapest GPU for running local LLMs?

It depends on how much VRAM you need. At 16GB the cheapest is currently the RX 7800 XT and at 24GB the RX 7900 XTX, both AMD and recommended only for llama.cpp / Ollama / LM Studio users; if you want NVIDIA, the 24GB pick is the RTX 3090. Live prices are in the tables.

Why not cheap data-center cards like the Tesla V100 or P40?

The V100 has no BF16 or INT8 tensor cores, so many kernels for new models fall back to slow FP16 paths; it is passively cooled, PCIe 3.0, and driver support can end at any time. The P40 is older still, with useful compute only in FP32. In 2026 their price per GB does not buy back the time you spend on them.

How do I buy a used GPU on eBay safely?

Choose Top Rated sellers with 30-day returns; read the listing for 'for parts / not working'; on arrival run memtest_vulkan and 30 minutes at full load. Import duties and shipping vary by country and are not in this site's prices.

Is it worth buying a modded 48GB RTX 4090?

Only if you accept zero warranty. The performance and memory are real, and it is the cheapest single 48GB card; the risk is the lifetime of the reflowed memory chips and the origin of the BIOS. A Radeon PRO W7900 or two 3090s are the safer routes to 48GB.

GPUs in this guide

Models in this guide