On a tight budget, card shopping needs one number: price per GB of VRAM, the used price divided by capacity. It compares across generations, brands, consumer and workstation cards alike. As of September 5, 2026 the cheapest card per tier is: RX 7800 XT at 16GB, RX 7900 XTX at 24GB, Tesla V100 32GB at 32GB (not recommended, see below), Radeon PRO W7900 at 48GB, the CMP 170HX mod and Instinct MI210 at 64GB, and the Mac Studio M4 Max at 128GB. The table updates live; the rest of this guide covers which cheap cards to skip and how to buy used.
Best Budget GPU for Local LLM Inference in 2026: The Cheapest Card at Every VRAM Tier
Translated from the 简体中文 edition · Read the original
Figures in the text and FAQ are as of Sep 5, 2026. Tables use the latest published prices; collection dates may differ by device.

Budget cards come down to one number: price per GB
Price per GB defines "cheap" precisely: how many gigabytes of VRAM a dollar buys, and therefore which model tier you can run. Two conditions apply:
- Bandwidth cannot be too low: VRAM decides whether a model runs, bandwidth decides how fast. Below 500 GB/s, models above 14B feel sluggish, so 256 to 273 GB/s machines like the DGX Spark and Strix Halo are excluded from this ranking.
- The ecosystem must work: CUDA is painless, ROCm works on RDNA 3 / 4 (see the state of AMD), oneAPI (Intel Arc) runs llama.cpp with thin documentation, and old architectures (Volta, Pascal) are losing software support.
The home page ranking is already sorted by price per GB; this guide slices it by tier.
The cheapest card at every VRAM tier
The cheapest and second-cheapest card in six tiers, ranked live by current used median:
AmazonJP
| VRAM tier | Cheapest card | AmazonJP | Price per GB | Bandwidth GB/s | Next cheapest | AmazonJP |
|---|---|---|---|---|---|---|
| 16 GB | $659 | $41.2 | 900 | $884 | ||
| 24 GB | $1,164 | $48.5 | 456 | $1,738 | ||
| 32 GB | $1,141 | $35.7 | 900 | $2,074 | ||
| 48 GB | No priced card in this tier | — | — | — | — | — |
| 64 GB | No priced card in this tier | — | — | — | — | — |
| 128 GB | $3,913 | $30.6 | 256 | $8,223 |
A tier covers cards from that capacity up to the next tier. Only cards with a current price are ranked.
Tier by tier:
- 16GB: the RX 7800 XT has the lowest price per GB; its 624 GB/s runs an 8B model near 57 tokens/s, which is fine. For CUDA look at a used RTX 4080. The catalog now covers the 22GB modified RTX 2080 Ti in the next capacity tier.
- 24GB: the RX 7900 XTX versus RTX 3090 question, see the state of AMD and the 3090 guide. This is the budget sweet spot: it runs 32B at 4-bit.
- 32GB: the Tesla V100 32GB is astonishingly cheap per GB, and the next section explains why to skip it; the real choices are the Radeon AI PRO R9700, the Arc Pro B70, or stretching to an RTX 5090.
- 48GB: the Radeon PRO W7900 has a warranty, two slots and 864 GB/s, steadier than a modded 48GB 4090; two 3090s are the other route, see multi-GPU basics.
- 64GB: the CMP 170HX 64GB mod (HBM2e, 1.5 TB/s, PCIe x4) and the Instinct MI210 are enthusiast options, passively cooled with no display outputs.
- 128GB: the Mac Studio M4 Max is the only 128GB device that works out of the box, see Mac unified memory.
Cheap cards to avoid
Three kinds of card are not worth it at any price per GB: old-architecture data-center cards, passively cooled accelerators, and modded cards without warranty.
AmazonJP
| Model | VRAM GB | Bandwidth GB/s | INT8 TOPS | Power W | AmazonJP | Price per GB |
|---|---|---|---|---|---|---|
| 32 | 900 | — | 250 | $1,141 | $35.7 | |
| 16 | 900 | — | 250 | $659 | $41.2 | |
| 22 | 616 | 215 | 250 | — | — | |
| 32 | 1,229 | 92 | 300 | — | — | |
| 40 | 1,560 | 404 | 250 | — | — | |
| 64 | 1,493 | 404 | 250 | — | — | |
| 48 | 1,008 | 661 | 450 | — | — | |
| 32 | 736 | 418 | 320 | — | — |
- Tesla V100 16GB / 32GB: Volta has FP16 tensor cores only, no BF16 or INT8, so many of llama.cpp's newer kernels take the slow path; PCIe 3.0; passive cooling that needs your own ducting; CUDA 13 dropped Volta, pinning you to old drivers. 900 GB/s of HBM2 looks great on paper for a 2017 card.
- RTX 2080 Ti 22GB (modified): upgraded from the original 11GB. Turing lacks native BF16. Memory capacity increases, but bandwidth and architecture remain unchanged; check memory stability and the seller’s support for the modification.
- Instinct MI100: first-generation CDNA, ROCm support in maintenance, half-rate BF16.
- CMP 170HX mods: a mining die with reflowed HBM2e; PCIe 1.1 x4 makes model loading take minutes, no display output, no warranty, nobody to repair it.
- RTX 4090 48GB / 4080 SUPER 32GB mods: real performance and real memory; the risk is the lifetime of the reflowed chips and where the BIOS came from, with no warranty at all. If you accept that, they are the cheapest single 48GB card.

How to buy used without getting burned
On eBay, lean on Top Rated sellers and 30-day returns; on Xianyu, on the platform inspection service and recorded tests; on both, retest yourself.
eBay:
- Choose Top Rated Plus sellers with returns accepted for 30 days; for private sellers, check the feedback history for previous GPU sales.
- Skip any listing whose title or description says for parts, as-is, artifacts or no display.
- On arrival check
nvidia-smi -qfor a power limit and BIOS version that match the model, then run memtest_vulkan. - Import duties and shipping are not in this site's prices; add them for your country.
Xianyu (for readers buying in China):
- Only sellers who offer the platform inspection service, which checks the card before releasing payment.
- Ask for a 30-minute
memtest_vulkanor OCCT VRAM test on video with the day's date visible. - Treat "mining bundle", "no returns" and other low-priced listings as mining cards and expect another 20% off before considering them.
- Run your own 30-minute full-load test within 24 hours of delivery and file a return at the first artifact.
Verdict: what to buy at each tier
One answer per tier: RX 7800 XT or a used RTX 4080 at 16GB, RX 7900 XTX or RTX 3090 at 24GB, Radeon PRO W7900 or two 3090s at 48GB, Mac Studio M4 Max at 128GB; the 32GB and 64GB tiers have no painless cheap card.
- The first budget priority is reaching 24GB: it is the threshold for 32B models and the sweet spot for price per GB.
- Do not buy old architectures or modded cards for the unit price unless you can repair them yourself.
- Run the three-year math before buying, see what running LLMs locally really costs.
FAQ
What is the cheapest GPU for running local LLMs?
It depends on how much VRAM you need. At 16GB the cheapest is currently the RX 7800 XT and at 24GB the RX 7900 XTX, both AMD and recommended only for llama.cpp / Ollama / LM Studio users; if you want NVIDIA, the 24GB pick is the RTX 3090. Live prices are in the tables.
Why not cheap data-center cards like the Tesla V100 or P40?
The V100 has no BF16 or INT8 tensor cores, so many kernels for new models fall back to slow FP16 paths; it is passively cooled, PCIe 3.0, and driver support can end at any time. The P40 is older still, with useful compute only in FP32. In 2026 their price per GB does not buy back the time you spend on them.
How do I buy a used GPU on eBay safely?
Choose Top Rated sellers with 30-day returns; read the listing for 'for parts / not working'; on arrival run memtest_vulkan and 30 minutes at full load. Import duties and shipping vary by country and are not in this site's prices.
Is it worth buying a modded 48GB RTX 4090?
Only if you accept zero warranty. The performance and memory are real, and it is the cheapest single 48GB card; the risk is the lifetime of the reflowed memory chips and the origin of the BIOS. A Radeon PRO W7900 or two 3090s are the safer routes to 48GB.