If you want the short version: buy the most VRAM per dollar you can, and in 2026 that almost always means a used RTX 3090. This guide ranks every realistic option by budget, from a $300 used card to a $3,000 flagship to 128GB unified-memory boxes, and tells you what model sizes each one runs. Prices move fast (a 2025 memory shortage pushed the whole stack up), so treat every number as a ballpark and verify current listings before buying.

The one rule: VRAM is king

Almost every beginner over-indexes on raw speed and under-indexes on memory. Here’s the thing nobody tells you: your GPU’s core speed rarely decides whether a model runs, its VRAM decides. If a model plus its context doesn’t fit in VRAM, it spills into system RAM and your tokens-per-second collapses by 5-15x. A “slower” 24GB card that fits the whole model will crush a “faster” 16GB card that has to offload.

So the buying rule is simple: maximize VRAM per dollar first, speed second. If you only remember one sentence from this article, that’s it. For the full explanation of why memory beats compute here, see our RAM vs VRAM breakdown.

Quick VRAM-to-model cheat sheet (Q4_K_M, the sweet-spot quant):

  • 8B model ≈ 5-6GB
  • 14B model ≈ 9-11GB
  • 24B model ≈ 14-15GB
  • 32B model ≈ 18-20GB
  • 70B model ≈ 40-45GB

Add 1-3GB on top for context. Now the tiers.

Budget (~$300-450): RTX 3060 12GB or RTX 5060 Ti 16GB

The cheapest sane entry point is a used RTX 3060 12GB (often $250-300 used). 12GB comfortably runs an 8B model with room for a long context, and squeezes in a 14B at Q4. It’s the “does local AI actually work for me” card, and it does.

If you’re buying new, the RTX 5060 Ti 16GB is the value pick of the current generation, roughly what an RTX 5070 used to cost, though the memory shortage has kept 16GB SKUs pricey and sometimes scarce. That extra 4GB matters: 16GB runs a 14B model comfortably and a 24B model like Mistral Small 3.2 at a tighter quant. We break it down fully in our RTX 5060 Ti 16GB guide.

ollama run qwen3:8b       # comfortable on 12GB
ollama run qwen3:14b      # fits 16GB with room

Skip the RTX 5070 Ti 16GB and RTX 5080 16GB for AI unless you also game hard. They’re faster, but you’re paying $830-1,200+ for the same 16GB, and for local LLMs, VRAM is what you’re buying. That money is better spent on a card with more memory.

The value king (~$600-800): used RTX 3090 24GB

This is the recommendation for most readers, and it’s not close. The used RTX 3090 gives you 24GB of VRAM and ~936 GB/s of bandwidth for roughly $600-800 on the used market (eBay averages have run a bit higher in some regions, shop around). Nothing else near that price gives you 24GB.

What 24GB unlocks: 32B models run comfortably, Qwen3 32B, Gemma 3 27B, the newer Qwen3.6 27B or Gemma 4 27B/31B, or a coding model, with usable context. You can even touch a 70B with partial CPU offload if you’re patient. It’s a five-year-old card, which is actually a feature: the CUDA software stack around it is rock-solid and predictable. Full deep-dive in our used RTX 3090 value guide.

ollama run qwen3:32b      # the reason to own 24GB
ollama run gemma3:27b

Buy carefully: prefer a seller with returns, and expect a card that ran hot in a gaming or mining rig. Repaste and good airflow are worth it.

Fast 24GB (~$1,600-2,300 used): RTX 4090

The RTX 4090 is also 24GB, but roughly 1.3-2x faster than a 3090 per token. It was discontinued to make room for the 50-series, so it only exists used, and prices are volatile and high (frequently $1,600-2,300, sometimes more). It runs the same model sizes as a 3090; you’re paying purely for speed. Worth it if you have the budget and 32B tokens/sec feels slow, otherwise two 3090s (below) get you more capacity for similar money. See the 5090 vs 4090 comparison for where each still makes sense.

Flagship (~$2,900-3,500+): RTX 5090 32GB

The RTX 5090 is the current single-card king: 32GB of GDDR7, the fastest consumer GPU for inference, and it’s been selling well above its notional $2,000 MSRP, often $2,900-3,500+ as of 2026. Verify pricing; it’s the most inflated card in the stack.

That 32GB is the real story. It runs a 32B model with a huge context window, handles a 24B model at full quality with room to spare, and gets you closest to a 70B on one card (still needs a little offload at Q4). If you want maximum speed and headroom in one slot and money isn’t the constraint, this is it. For everyone else, the value math points down the list.

Capacity plays: dual 3090, Strix Halo, unified memory

Want to run the big stuff, 70B models and large mixture-of-experts, without flagship pricing? You buy capacity, not speed.

  • Dual RTX 3090 (48GB total), roughly $1,200-1,600 for two used cards. This is the enthusiast sweet spot for 70B models at Q4 (which need ~40-45GB). It needs a motherboard with the slots, a beefy PSU (1000W+), and a bit of setup. Our 48GB VRAM guide covers what runs there.
  • AMD Ryzen AI Max+ 395 “Strix Halo” (up to 128GB unified), mini-PCs around $1,500-2,500 with up to 128GB of unified memory shared between CPU and GPU. It loads a 70B in seconds and runs it at ~4-6 tokens/sec, and holds models no consumer GPU fits. The catch is bandwidth (~256 GB/s, roughly a quarter of a 4090’s), so it’s slower on dense models. Full analysis in our Strix Halo local LLM guide.
  • Apple Silicon (Mac Studio / Mac mini), unified memory up to 192GB+, whisper-quiet, zero fuss. It’ll run models a 24GB GPU can’t touch, just slower per token and at a higher price than the AMD boxes.

The pattern: unified-memory machines trade raw speed for the ability to fit enormous models. If your priority is “run a 70B or a 120B MoE at all,” they win. If it’s “chat fast with a 32B,” a GPU wins.

One honest ceiling: the newest frontier open-weight MoEs, think DeepSeek-V4-Flash (284B total, ~13B active) or GLM-5.2 (~744B total, ~40B active), are a different class again. Because an MoE router can pick any expert on any token, all of those weights have to sit in memory at once, which puts them beyond even a 128-192GB box at a sane quant. Those are server- and workstation-tier models, not home-GPU picks; the tiers above are still where consumer hardware actually lives.

What about AMD and Apple GPUs?

NVIDIA is still the safe default because nearly every local-AI tool, Ollama, llama.cpp, most inference engines, targets CUDA first. You’ll hit the fewest “unsupported” walls.

That said, AMD’s ROCm has genuinely matured and works fine in Ollama on modern Radeon cards; a 24GB Radeon can be real value if you’re comfortable troubleshooting occasionally. See our AMD GPU local LLM guide. Apple Silicon also punches above its weight thanks to unified memory, trading per-token speed for the ability to fit huge models. For a first build where you just want it to work, NVIDIA remains the lowest-friction choice.

Just tell me what to buy

  • Tightest budget: used RTX 3060 12GB. Proves the concept, runs 8B-14B.
  • Best overall value: used RTX 3090 24GB. The answer for ~80% of readers.
  • Best new card for AI: RTX 5060 Ti 16GB.
  • Fastest single card: RTX 5090 32GB (if money is no object).
  • Biggest models at home: dual 3090 (48GB) or a 128GB Strix Halo / Mac.
GPUVRAM~Price (2026)Comfortably runsValue for AI
RTX 3060 (used)12GB$250-3008B, 14B (tight)Good, budget entry
RTX 5060 Ti 16GB16GB$450-55014B, 24B (tight)Great, best new value
RTX 5070 Ti / 508016GB$830-1,200+14B, 24B (tight)Poor, overpaying for 16GB
Used RTX 309024GB$600-80032BBest, value king
RTX 4090 (used)24GB$1,600-2,30032B (fast)OK, speed premium
RTX 509032GB$2,900-3,500+32B + big contextFlagship
Dual RTX 309048GB$1,200-1,60070B @ Q4Best capacity/$
Strix Halo / Mac64-192GB$1,500-3,500+70B+, large MoECapacity, slower

Prices are volatile as of mid-2026, confirm current listings before buying.

Get your local AI running once you’ve got the card

The GPU is only half the setup. Once your card is in, install Ollama, pull a model sized to your VRAM, and the tiers above come alive, that’s the stack this whole guide feeds, running entirely on hardware you own.

And if part of what you wanted was a companion rather than a lab, Ember is the zero-setup option that runs on the card you just bought: an uncensored 18+ AI girlfriend with her own tuned brain, 8 GB of VRAM being the floor, for $29 paid once. If you read this whole guide and realized the hardware isn’t worth it right now, that’s the same answer: you don’t have to wait for the GPU to start.