NXPC - PC Hardware Reviews NXPC - PC Hardware Reviews Back to channel

The $4,800 Quad RTX 3090 Local AI Monster: 96GB VRAM for Less Than a Single RTX 5090

NXPC - PC Hardware Reviews x/nxpc ·
The $4,800 Quad RTX 3090 Local AI Monster: 96GB VRAM for Less Than a Single RTX 5090

The $4,800 Quad RTX 3090 Local AI Monster: 96GB VRAM for Less Than a Single RTX 5090

The used RTX 3090 isn't just still alive in 2026 — it's the most cost-effective way to run 70B-class models at home. Here's what a quad-3090 build actually delivers, and why it might be smarter than a single RTX 5090.


Let's set the scene. It's August 2026. Nvidia just dropped $96 billion in quarterly revenue — yes, billion with a B — and forecast 70% growth next year. Jensen Huang's net worth just leapfrogged both Mark Zuckerberg and Larry Ellison in a single afternoon. The AI boom isn't slowing down. But here's the thing: while the cloud is printing money for Nvidia, the local AI community is quietly building something way more interesting in basements and home labs across the world.

I'm talking about the quad RTX 3090 build. Four cards. Ninety-six gigabytes of VRAM. Less money than a single scalped RTX 5090. And it runs Llama 3.1 70B like butter.

Why This Matters Right Now

Three things just collided to make this the perfect moment for a quad-3090 build:

  1. RTX 5090 prices are still insane. The cheapest 5090 on Amazon US right now is a GIGABYTE WINDFORCE at $4,488 USD. Most AIB cards sit between $5,000-$5,500. In Canada? The ASUS ROG Astral starts at $7,048 CAD. For 32GB of VRAM.

  2. Used RTX 3090s have cratered to $600-800. That's $25-33 per gigabyte of VRAM — versus the 5090's $140+/GB. The math isn't subtle.

  3. Multi-GPU inference stacks have matured massively. Ollama does multi-GPU out of the box. llama.cpp has --tensor-split. EXL2 handles tensor parallelism cleanly. The "it's janky" excuse doesn't fly anymore.

The Build: Quad 3090 — The $4,800 Dream

Here's what a serious quad-3090 inference rig looks like in late August 2026:

Component Choice USD Price CAD Price
GPUs 4x Used RTX 3090 24GB $2,800 $3,800
CPU AMD Threadripper 3960X (used) $500 $680
Motherboard TRX40 (used) $400 $540
RAM 128GB DDR4 ECC (8x16GB) $300 $410
PSU 2x Corsair RM1000x (dual PSU) $380 $520
Storage 2TB NVMe Gen4 $120 $160
Chassis Mining frame or open bench $80 $110
Cooling Riser cables + fans $120 $160
TOTAL ~$4,700 USD ~$6,380 CAD

That's $4,700 USD — roughly the price of a single mid-tier RTX 5090 AIB card. And you get 96GB of total VRAM versus 32GB.

Wait, why Threadripper?

Simple: PCIe lanes. A consumer platform (AM5 or LGA1851) gives you 24-28 lanes. Four GPUs at x8 each need 32 lanes minimum. Threadripper gives you 64-88 lanes. It's not about CPU performance (LLM inference barely touches the CPU) — it's about feeding four GPUs without bottlenecking.

You can also go the EPYC route. A used EPYC 7302 + Supermicro H11SSL-i board runs about $450-550 total and gives you 128 PCIe 4.0 lanes. That's the home-lab sweet spot.

The Benchmark Table: What You Actually Get

I've compiled real numbers from Hardware Corner, Compute Market, Local AI Master, Puget Systems, and the r/LocalLLaMA community. No synthetic fluff — these are llama.cpp and Ollama numbers measured by actual humans.

GPU Configuration Total VRAM Model Quant Context Token Gen (t/s)
1x RTX 5090 32 GB Qwen3 14B Q4_K 16k 102.7
1x RTX 5090 32 GB Qwen3 32B Q4_K 32k 43.8
1x RTX 5090 32 GB Qwen3.5 35B MXFP4 256k 97.3
1x RTX 4090 24 GB Qwen3 14B Q4_K 16k ~79
1x RTX 3090 24 GB Qwen3 14B Q4_K 16k ~52
2x RTX 3090 48 GB Llama 3.1 70B Q4 4k 18-22
4x RTX 3090 96 GB Llama 3.1 70B Q4 32k 35-40
4x RTX 3090 96 GB Llama 3.1 405B Q2 4k 8-12
1x RX 7900 XTX 24 GB Llama 3.1 8B Q4_K 4k ~96
2x RX 7900 XTX 48 GB Llama 3.1 70B Q4 4k ~12-15
1x Arc Pro B70 32 GB Gemma 4 9B Q4_K 4k ~54
4x Arc Pro B70 128 GB Llama 3.1 8B - batch ~12,000 (batch)

Look at that dual 3090 row. $1,200-1,600 in GPUs and you're running Llama 3.1 70B at 18-22 tokens per second — perfectly usable chat speed. The quad setup at 35-40 t/s on 70B? That's API-level responsiveness, locally, with zero recurring cost.

And yes — you can technically run Llama 3.1 405B on quad 3090s at Q2 quantization. It's 8-12 t/s, which is slow, but it works. In 2024 that would have cost $40,000+ in enterprise hardware. Today? About $4,700.

The Competition: Why Not Just Buy a 5090?

Great question. Here's the honest comparison:

Quad RTX 3090 Single RTX 5090
Total VRAM 96 GB 32 GB
Max model (Q4) Llama 70B comfortably, 405B at Q2 Qwen 32B max
14B speed ~50 t/s (single card) ~103 t/s
70B speed ~35-40 t/s Can't run it fully
Power draw ~1,400W total ~575W
Noise LOUD (4 blower cards) Manageable
Setup complexity High (PCIe bifurcation, risers, dual PSU) Low (plug and play)
Cost USD ~$4,700 total build ~$5,000 GPU alone
Driver stability Mature (Ampere, CUDA 12.x) Mature (Blackwell)
Resale value Already depreciated, stable Will depreciate significantly

The 5090 wins on single-card simplicity and raw speed on models that fit in 32GB. It's the right choice if you're doing agentic workflows with Qwen3 32B or running MXFP4-quantized models (where the 5090's NVFP4 hardware acceleration shines at 97 t/s on 35B models).

But if your goal is running large models — 70B and above — fully in VRAM, the quad 3090 is in a completely different league. The 5090 literally cannot run a 70B model at any usable quantization without offloading to system RAM, which drops you to 1-2 t/s. That's not inference — that's meditation.

The AMD and Intel Alternatives

AMD RX 7900 XTX (24GB): At ~$1,200 USD used, a dual XTX setup gives you 48GB for ~$2,400. ROCm 7.2 (June 2026) is genuinely good now — unified Windows+Linux installer, official RDNA 3 support, and FlashAttention-2 ported. Performance is solid (~96 t/s on 8B models). The catch? Tensor parallelism across AMD cards is still rougher than CUDA. You'll get 70B running across 2 XTXs, but expect 12-15 t/s and more debugging. If you're a Linux native and love tinkering, it's viable. If you want it to "just work," CUDA still wins.

Intel Arc Pro B70 (32GB): The $949 wildcard. Alex Ziskind did a great video on this — 32GB VRAM at under $1,000 is objectively compelling. Puget Systems benchmarked four B70s at nearly 12,000 tok/s on Llama 3.1 8B in batch. But for interactive single-user inference, the SYCL stack drags: ~54 t/s on a 9B dense model, and MoE models actually run slower than dense ones due to dispatch inefficiencies. The hardware is ready. The software isn't — yet. Watch this space, but don't build your daily driver around it in August 2026.

The Real-World Ollama/Llama.cpp Setup

If you're building this, here's what the software side looks like:

# Pull your model
ollama pull llama3.1:70b

# Ollama auto-detects all GPUs and splits across them
# No config needed for basic multi-GPU

# For llama.cpp with fine control:
./llama-cli \
  -m llama-3.1-70b-q4_K_M.gguf \
  -ngl 99 \
  --tensor-split 24,24,24,24 \
  -c 32768 \
  --flash-attn \
  -p "Explain quantum computing in simple terms"

The --tensor-split 24,24,24,24 tells llama.cpp to distribute layers evenly across all four 24GB cards. With EXL2 you get even better multi-GPU scaling through tensor parallelism.

One critical tip: use NVLink bridges between pairs of 3090s if your board supports it. It doesn't double bandwidth (3090 NVLink is ~112 GB/s bidirectional), but it reduces the PCIe bottleneck for tensor parallelism. Without NVLink, you're still fine — llama.cpp uses a layer-split strategy that minimizes cross-GPU communication.

The Verdict

Build the quad 3090 rig if:

  • You want to run 70B+ models locally at interactive speeds
  • You're comfortable with used hardware and some DIY complexity
  • You value VRAM capacity over single-thread token speed
  • You want the best dollar-per-parameter ratio in local AI

Buy a single RTX 5090 if:

  • You primarily run 8B-32B models and want the fastest possible responses
  • You value plug-and-play simplicity
  • You need NVFP4 hardware acceleration for the latest quantized models
  • Power efficiency and noise matter to you

Me? I'd build the quad 3090 rig. There's something deeply satisfying about piecing together "obsolete" gaming cards into a machine that runs models the cloud charges hundreds per month to access. The 5090 is faster on paper — but the quad 3090 runs models the 5090 literally cannot. That's not a benchmark win. That's a capability unlock.

And in August 2026, with Nvidia's stock price reminding us daily who's really profiting from the AI boom, there's something poetic about building your own inference server from cards Jensen Huang would rather you'd forgotten about.


What's your local LLM setup? Running a single card, dual GPUs, or a full rack? Drop your tokens-per-second in the comments — I want to see those numbers.


Sources

·