What is the best GP...
 
Notifications
Clear all

What is the best GPU for running DeepSeek locally?

3 Posts
4 Users
0 Reactions
193 Views
0
Topic starter

So Ive been running local models for a while now and usually my dual 3060 setup handles most things fine but man this DeepSeek stuff is a different beast entirely. Im trying to get the full DeepSeek-V3 or even just a decent quant of R1 running without it taking ten years to generate a single line of code. Every time I try to load anything bigger than the distilled 7b or 14b versions I just hit a wall with OOM errors or the tokens per second are just abysmal. Its honestly so frustrating because I really need the reasoning capabilities for this backend project I'm finishing up by Friday but my hardware is just choking. I have about 1800 bucks to drop on a new card or two right now and I need to buy it like today or tomorrow so I can actually get some work done.

  • Must have at least 24gb vram because anything less feels like a waste for these MoE models
  • Good support for FP8 or whatever the latest quantization method is for DeepSeek
  • Needs to fit in a standard mid-tower case because I dont have room for a massive server rack
  • Looking for the best bang for buck in terms of inference speed

Should I just go for a used 3090 or is it worth stretching for a 4090 or maybe even some weird multi-GPU setup with 4060 tis? I keep seeing conflicting stuff about memory bandwidth being the main bottleneck for these models and I dont want to waste money on something that wont actually fix the lag...


3 Answers
12

I totally get your frustration with those awful OOM errors, but you are going to love the jump to a single massive card! Honestly, as someone who prefers to keep things super safe and reliable without messing with crazy configurations, I highly recommend staying away from the multi-GPU 4060 ti idea. It sounds like a total nightmare to set up and get working correctly with DeepSeek, especially when you are on a tight deadline. You should absolutely look at getting a brand new ASUS TUF Gaming GeForce RTX 4090 24GB GDDR6X if you can find one close to your budget, or even a certified refurbished EVGA GeForce RTX 3090 FTW3 Ultra Gaming 24GB from a reputable dealer to keep that warranty safety net! Personally, I went with the 4090 route and the FP8 support is just fantastic! It runs the quantized models so fast, it completely blew my mind compared to my old setup. The memory bandwidth on the 4090 is incredible, and it fits right into a standard mid-tower without any weird server rack modifications. Going with a single powerful card is just so much safer for your system stability, especially with your Friday deadline coming up so fast! Let me know if you need help picking out a specific model that fits your power supply, I am more than happy to double-check the specs for you so you dont run into any compatibility issues!


12

Snag two used NVIDIA GeForce RTX 3090 Founders Edition 24GB units for $1500 total.

  • 48GB VRAM pool
  • Saves cash for cooling Honestly, VRAM capacity is king for R1. A single card wont cut it.


3

> stretching for a 4090 Just saw this. The NVIDIA GeForce RTX 4090 24GB bandwidth is king for FP8. What is your PSU wattage? You'll need at least 850W.


Share: