im trying to get deepseek running locally on my own machine but honestly im kinda drowning in all the different advice online. i keep seeing people argue about whether you really need massive vram or if you can get away with just a decent mac studio setup and some people are swearing by dual 3090s while others say thats overkill unless youre doing training.
my use case is mostly just experimenting with local llms for my own research notes and maybe some light coding help since i dont want my data sitting on some random server. budget wise im hoping to stay under 2k maybe 2.5k if i really have to. i looked at some benchmarks for the 70b models and some folks say you need like 48gb of vram minimum but then i saw a reddit thread claiming you can run it fine with less if you use quantization. i just dont know which path to take. should i just build a pc with a used card or should i look for a high memory mac? any tips from people who have actually got it running smooth would be a lifesaver because i dont want to drop that much cash and realize i cant even run the models i want at a decent speed...
Regarding what #2 said about peace of mind, they hit the nail on the head. Dealing with multi-GPU setups on Windows or Linux can be a total nightmare if you just want to get your work done. Coming back to this, I honestly find that a solid single-card build is the sweet spot for most folks. I run my setup with a NVIDIA GeForce RTX 4090 24GB GDDR6X and it handles quantized 70b models perfectly fine for research and coding tasks. You dont need 48gb of VRAM if you use GGUF quantization. Just make sure you have enough system RAM to back it up. If you go the PC route, spend the extra cash on faster memory like Corsair Vengeance 64GB DDR5 6000MHz. It keeps everything responsive without the headache of balancing power draws across two cards.
Honestly, if you want peace of mind and reliability, macs are honestly hard to beat for this. I get the whole dual-GPU hype, but dealing with power supplies, PCIe lanes, and weird driver stuff when you just want to run research notes is a total headache. If you grab a Apple Mac Studio M2 Ultra 64GB Unified Memory, you get that unified memory which basically lets the GPU tap into your system RAM. It makes running 70b models with decent quantization way less of a nightmare than trying to bridge two cards.
Honestly, skip the Mac route unless you need the portability. In my experience, building a rig with dual used cards is the way to go. You get way better bang for your buck on memory bandwidth. Just grab any high-end NVIDIA GeForce card from the last few generations and you will be set. Honestly, once you have enough VRAM for quantization, everything just feels so much snappier for coding tasks.