AI Workload GPU Price Tracker
Consumer RTX, Workstation RTX A-Series & Datacenter Tesla / A100 / Instinct
For running local LLMs, Stable Diffusion, or any other AI workload, VRAM — not raw gaming benchmarks — is usually the number that decides what you can actually load. This tracks current $/GB VRAM pricing across the GPUs homelabbers actually buy for AI: consumer RTX 3090/4090 cards, workstation RTX A-series cards, and used datacenter compute cards like the Tesla P40/V100, NVIDIA L4/A100, and AMD Instinct MI50.
Why VRAM Is the Number That Matters
A GPU's shader count and memory bandwidth affect how fast a model runs, but VRAM decides whether the model fits at all. If a model doesn't fit in VRAM, it either won't load or has to spill over to much slower system RAM or disk offloading. That makes $/GB VRAM — not just raw price — the useful number to compare across wildly different card types.
This is also why older datacenter cards like the Tesla P40 stay popular in homelab AI circles years after their gaming-era equivalents would be considered obsolete: 24GB of VRAM still fits plenty of useful quantized models, and depreciated enterprise surplus pricing means that VRAM is often the cheapest on the used market of any card type here.
Quantization (running a model at reduced numeric precision, like 4-bit instead of 16-bit) is the other lever — it roughly quarters the VRAM a given model needs at some cost to output quality, which is why a 24GB card can run models that would need 80GB+ at full precision.
The Three GPU Classes on This Page
Consumer (RTX 3090 / 4090)
Usually the best raw price per GB of VRAM, since they're mass-produced gaming cards. No ECC memory, and warranty/duty-cycle expectations assume gaming use, not 24/7 compute.
Prosumer (RTX A-Series)
ECC memory, blower-style coolers built for dense multi-GPU chassis, and higher VRAM configurations than their consumer counterparts — at a real price premium.
Datacenter (Tesla / A100 / Instinct)
Purpose-built for sustained compute, often the cheapest $/GB VRAM on older generations as enterprise fleets retire them — but expect passive cooling that needs a server chassis or an aftermarket fan.
Loading…
| GPU | Brand | Class | Condition | VRAM | Price | $/GB VRAM | Deal |
|---|---|---|---|---|---|---|---|
AMD Instinct MI50 32GB HBM2 Datacenter GPU - Tested Working Instinct MI50 · eBay | AMD | Datacenter | Used | 32 GB | $179.00 | $5.59/GB | View Deal |
NVIDIA Tesla P40 24GB Datacenter GPU - Pulled from Server, Untested Tesla P40 · eBay | NVIDIA | Datacenter | Used | 24 GB | $139.00 | $5.79/GB | View Deal |
NVIDIA Tesla P40 24GB GDDR5 Datacenter GPU - Tested Tesla P40 · eBay | NVIDIA | Datacenter | Used | 24 GB | $169.00 | $7.04/GB | View Deal |
AMD Instinct MI50 16GB HBM2 Datacenter GPU Instinct MI50 · eBay | AMD | Datacenter | Used | 16 GB | $119.00 | $7.44/GB | View Deal |
NVIDIA Tesla V100 32GB PCIe Datacenter GPU Tesla V100 · eBay | NVIDIA | Datacenter | Used | 32 GB | $459.00 | $14.34/GB | View Deal |
NVIDIA Tesla V100 16GB SXM2 Datacenter GPU Tesla V100 · eBay | NVIDIA | Datacenter | Used | 16 GB | $289.00 | $18.06/GB | View Deal |
NVIDIA RTX A4000 16GB GDDR6 Workstation Graphics Card RTX A4000 · eBay | NVIDIA | Prosumer | Refurbished | 16 GB | $399.00 | $24.94/GB | View Deal |
NVIDIA RTX A5000 24GB GDDR6 Workstation Graphics Card RTX A5000 · eBay | NVIDIA | Prosumer | Used | 24 GB | $649.00 | $27.04/GB | View Deal |
NVIDIA GeForce RTX 3090 24GB Founders Edition GDDR6X Graphics Card RTX 3090 · eBay | NVIDIA | Consumer | Used | 24 GB | $679.00 | $28.29/GB | View Deal |
NVIDIA GeForce RTX 3090 Ti 24GB Graphics Card - Tested Working RTX 3090 Ti · eBay | NVIDIA | Consumer | Used | 24 GB | $799.00 | $33.29/GB | View Deal |
NVIDIA RTX A6000 48GB GDDR6 Workstation Graphics Card RTX A6000 · eBay | NVIDIA | Prosumer | Used | 48 GB | $1,949.00 | $40.60/GB | View Deal |
NVIDIA GeForce RTX 4090 24GB Graphics Card RTX 4090 · eBay | NVIDIA | Consumer | Used | 24 GB | $1,599.00 | $66.63/GB | View Deal |
NVIDIA RTX 4090 24GB Founders Edition - New Open Box RTX 4090 · eBay | NVIDIA | Consumer | New | 24 GB | $1,899.00 | $79.13/GB | View Deal |
NVIDIA L4 24GB Tensor Core GPU - New L4 · eBay | NVIDIA | Datacenter | New | 24 GB | $2,399.00 | $99.96/GB | View Deal |
NVIDIA A100 40GB PCIe Datacenter GPU A100 · eBay | NVIDIA | Datacenter | Used | 40 GB | $4,899.00 | $122.48/GB | View Deal |
No GPUs match those filters. Try widening your search.
Frequently Asked Questions
How much VRAM do I need to run a local LLM?
As a rough rule of thumb, a model at full FP16 precision needs about 2GB of VRAM per billion parameters, while a 4-bit quantized version needs roughly 0.5-0.7GB per billion parameters. That puts a quantized 7-8B model at 8-12GB, a 13B model at around 16GB, a 30-34B model at 24GB, and a 70B model at 48GB or more. Larger frontier-scale models generally require 80GB+ or splitting the model across multiple GPUs.
Do Tesla and other datacenter GPUs need extra cooling to use outside a server?
Usually yes. Cards like the Tesla P40, P100, and V100 (PCIe versions) are passively cooled and rely on the high static-pressure airflow inside a server chassis to stay cool. Running one in a desktop case or open-air rig typically requires an aftermarket blower fan or 3D-printed shroud, or it will overheat and throttle.
Is a used mining GPU safe to buy for AI compute?
Generally yes. Mining is a steady, moderate 24/7 load that's arguably gentler on a card than the thermal cycling of gaming, and the VRAM chips that matter most for AI work see minimal wear either way. Watch listings for mentions of reflowed or reballed memory, and buy from sellers with a return policy so you can verify the card before the window closes.
What's the difference between consumer, prosumer, and datacenter GPUs for AI?
Consumer cards (RTX 3090/4090) offer the best raw price per GB of VRAM but are built for gaming, without ECC memory or long-duty-cycle warranty terms. Prosumer/workstation cards (RTX A-series) add ECC memory, blower coolers suited to dense multi-GPU builds, and higher VRAM configurations at a price premium. Datacenter cards (Tesla, A100, L4, Instinct) are purpose-built for sustained compute and often have the lowest $/GB VRAM on older generations, but need server-style airflow and sometimes different power connectors.
Can I combine multiple GPUs to run a bigger model?
Yes. Tools like llama.cpp, vLLM, and Ollama support splitting a model's layers across multiple GPUs' VRAM. Most post-30-series consumer cards no longer support NVLink, so multi-GPU setups communicate over PCIe instead, which is slower than NVLink but works fine for inference. Prosumer and datacenter cards more often retain NVLink or NVSwitch support for tighter multi-GPU scaling.