You’re mid-purchase and the question is simple: does the RTX 5090’s 32GB of VRAM and roughly 1.79 TB/s of bandwidth justify paying nearly double a used 4090? Short version, the 5090 is the fastest consumer card you can put a local model on, but the 8GB jump from 24 to 32GB does not unlock the tier most people are dreaming about (a 70B living entirely in VRAM). Here’s exactly what the extra memory and bandwidth buy you, where a used card is the smarter play, and a verdict per use case.

The spec sheet that actually matters

Ignore ray-tracing and gaming FPS. For running LLMs, three numbers decide everything: VRAM (how big a model fits), memory bandwidth (how fast it generates tokens), and price (whether the math makes sense).

SpecRTX 4090RTX 5090
VRAM24GB GDDR6X32GB GDDR7
Memory bandwidth~1.0 TB/s~1.79 TB/s
Memory bus384-bit512-bit
TDP450W575W
Launch MSRP$1,599 (discontinued)$1,999 (Jan 2025)
Street price, 2026~$1,600-2,500 (mostly used)~$2,000-3,200
Biggest comfy model (Q4)32B with modest context32B at Q6-Q8 or big context
~tok/s, 8B Q4 (single stream)~140-160~200-215
~tok/s, 70B Q4 (with offload)~35~50-55

Prices swing week to week, verify current listings before you buy. The 4090’s number is weird because NVIDIA discontinued it in late 2024, so it now floats between a cheap used unit and an absurd new-in-box scalp.

What the extra 8GB (24 → 32GB) really unlocks

This is the crux, so let’s be precise. VRAM sets the ceiling on model size and context. Rough Q4_K_M footprints: a ~7-8B model is 5-6GB, a ~14B is 9-11GB, a ~24B is 14-15GB, a ~32B is 18-20GB, and a ~70B is 40-45GB.

On 24GB, a 32B at Q4 fits with room for a normal context window, that’s already a strong tier, covered in the best models for 24GB VRAM. What the 5090’s extra 8GB adds is headroom, not a new class of model:

  • Higher-quality quants. Run that same 32B at Q6 or Q8 instead of Q4, noticeably crisper reasoning and less quant degradation.
  • Much longer context. The KV cache grows with context length and eats VRAM fast. 32GB lets you push a 24-32B model to 32K-64K tokens without spilling into slow system RAM.
  • Two models at once. Keep a 24B chat model resident and a vision or embedding model loaded, handy for RAG, OCR, or an always-on assistant that doesn’t reload on every switch.
  • Less offload, more speed. Anything that was borderline on 24GB (spilling a few layers into system RAM) now stays fully on the GPU, which is often a bigger real-world speedup than the raw bandwidth gain.

Bandwidth is why it feels faster

Here’s the thing nobody tells first-time buyers: token generation is memory-bandwidth-bound, not compute-bound. Every token, the GPU streams the model’s weights through its memory bus. More bandwidth = more tokens per second, almost linearly, for single-stream chat.

The 5090’s ~1.79 TB/s versus the 4090’s ~1.0 TB/s is a ~78% bandwidth jump. In practice you see roughly 25-35% more tokens per second on typical models, scaling up toward 40-50% on larger models where the card is purely bandwidth-limited. On an 8B model expect the 5090 in the low 200s of tok/s versus the 4090 around 145-160. Both are far past what you can read, see what tokens per second is actually usable, so for pure chat this delta is a luxury, not a fix.

The 70B trap

If your reason for eyeing the 5090 is “so I can finally run a 70B locally”, stop. A 70B at Q4 needs ~40-45GB. 32GB doesn’t fit it. You’d be running a heavily-quantized (Q3 or lower) version that degrades quality, or offloading layers to system RAM and dropping to ~50 tok/s at best.

For a 70B that lives entirely in VRAM at a decent quant, you want 48GB, which realistically means two 24GB cards. What it actually takes to run a 70B is a different budget conversation, and dual used 3090s often beat a single 5090 for that specific goal. The 5090 is the king of the 32B-and-below, but faster and roomier tier. Buy it for that, not for 70B fantasies.

When a used 4090 or 3090 is the smarter buy

Price is where the 5090’s story gets complicated. As of 2026, a 5090 runs roughly $2,000-3,200, a used 4090 sits around $1,600-2,500 (discontinued, so oddly expensive), and a used 3090, still 24GB, goes for roughly $800-1,100.

That makes the 3090 the value king it’s been for years. It runs the exact same models a 4090 does (both are 24GB), just slower thanks to lower bandwidth (~936 GB/s). If your workload is chat, roleplay, or a coding model up to 32B and you don’t need bleeding-edge speed, a used 3090 is hard to beat on dollars-per-capability.

The 4090 is the awkward one. Discontinued and scalped, it often costs nearly as much as a 5090 while offering less VRAM and far less bandwidth. Buy one only if you find a genuinely cheap used unit. Otherwise the choice collapses to: cheap 24GB (used 3090) or fast-and-roomy 32GB (5090), the 4090 is a trap in the middle. For the full ladder across every budget, see the best local LLM GPUs of 2026 by budget.

Verdict by use case

  • Chat / companion / roleplay: A 24GB card is plenty. A used 3090 or a fast 14-24B model already saturates readable speed. The 5090 is overkill here unless you also do other GPU work.
  • Coding assistant: The 5090’s headroom shines, run a 32B coder at a high quant with long context for whole-file work. A 24GB card handles it at Q4 with tighter context. See the best coding model for your VRAM.
  • Serious tinkerer / image + LLM / RAG: This is the 5090’s sweet spot. 32GB lets you juggle a big model, long context, and a second model (or Flux image gen) without reloading. Worth it.
  • 70B ambitions: Neither card fully solves this. Save for 48GB+ (dual 3090s) instead of assuming the 5090 gets you there.

Running your own AI on that new card

Whichever card you land on, the point of owning it is that the model runs on your silicon, nothing leaves the machine, no subscription, no logging. Pull an abliterated 32B, wire up a persona and memory, and every gigabyte of that 24 or 32GB is working for you. And if you also want an AI companion without the wiring, Ember is the packaged shortcut: an uncensored 18+ companion that takes roughly 7GB, a rounding error on either of these cards.

And if you crunch the numbers and a $2,000 GPU doesn’t make sense for how much you’d actually use it, that’s a completely fair call, a more modest card and a smaller quantized model still run a fully private local assistant, and the hosted companion route needs no GPU at all. The rig is a means, not the goal; pick the path that fits your budget and privacy needs, not the one with the biggest VRAM number.