What is the best GP...
 
Notifications
Clear all

What is the best GPU for running DeepSeek-V3 locally?

5 Posts
6 Users
0 Reactions
665 Views
0
Topic starter

Hey everyone! I'm trying to figure out which GPU can actually handle DeepSeek-V3 at home. I'm specifically worried about:

  • VRAM requirements for 4-bit quants
  • Multi-GPU scaling vs single pro cards

I want to stop paying for API calls for privacy reasons. What is the best GPU setup to run this smoothly?


5 Answers
12

Honestly, DeepSeek-V3 is a beast so you're gonna need massive VRAM. For a home setup, picking up used NVIDIA GeForce RTX 3090 24GB GDDR6X cards is usually the most cost-effective move. You get 24GB for way less than pro gear. Just a heads up, you'll need several of them in a multi-GPU rig to run a 4-bit quant smoothly, but its basically the best bang for your buck.


10

I stumbled upon this while looking for V3 benchmarks myself. Ive spent way too long trying to fit these huge MoE models into home builds lately. Tbh, the math for DeepSeek-V3 is pretty brutal. At 4-bit, youre looking at around 350GB to 400GB just for the weights and KV cache. Even with a bunch of consumer cards, youll run out of room fast. I actually started looking into NVIDIA RTX 6000 Ada Generation 48GB GDDR6 units for the better VRAM density, but the cost is definitely wild. Another route Ive seen some folks take is the Apple Mac Studio M2 Ultra 192GB Unified Memory, though youd be forced into a heavier quantization like 2-bit to fit it all. If you go the multi-GPU route on a PC, just make sure your motherboard like the ASUS Pro WS WRX90E-SAGE SE WIFI can actually handle all those PCIe lanes. Bottlenecking is a real vibe killer when youre waiting for tokens...


3

Facts.


3

Man, I spent the whole weekend fighting with my rig trying to get this thing to breathe. I finally ditched the idea of just piling up consumer cards because the power draw was literally melting my breaker! Seriously, my room felt like a sauna. Here is what I learned from my struggle:

  • The bottleneck isnt always the raw teraflops, its the bus speed between cards. If you go multi-GPU, make sure your motherboard has enough lanes or you'll be waiting forever for tokens.
  • I switched over to a workstation-grade board and it changed everything. The stability is just miles ahead of those consumer gaming setups that kept crashing whenever the context window got too big.
  • Honestly, stop obsessing over the newest flagship gear. I found that older enterprise equipment handles the memory bandwidth way better for these massive models. It feels amazing now that its finally running, but yeah, the learning curve is steep! You really gotta be patient with the cooling setups too.


1

Helpful thread 👍


Share: