I'm looking to set up a dedicated rig to run DeepSeek models locally for my coding projects. I'm specifically curious about the VRAM requirements for the 67B version versus the newer R1 models. Should I prioritize multiple 3090s or go for a high-end Mac? What's the most cost-effective hardware setup for smooth inference right now?
For your situation, I honestly think the hardware landscape for these big models is kinda frustrating right now. I've spent thousands trying to get smooth inference on the full R1, and unfortunately, it's just not as good as expected on consumer gear because of the sheer weight of the parameters.
Just saw this and honestly dont waste ur cash. I spent way too much on gear before realizing 4-bit quants are the way to go for R1.
I have been obsessing over this for the last few weeks while trying to figure out my own setup. Tbh the market is wild right now and it feels like there is a huge divide between the two main ecosystems.
Works great for me
> Should I prioritize multiple 3090s or go for a high-end Mac? What's the most cost-effective hardware setup for smooth inference right now? ^ This. Also, I have been dealing with this exact same dilemma for over a month now and honestly I am still no closer to a real answer. Its so frustrating trying to weigh the massive VRAM capacity of those unified memory systems against the raw speed of a multi-GPU build. Ngl every time I think I have finally made a choice I find a new benchmark that makes me second guess everything... I am basically just stuck in research hell right now.
@Reply #7 - good point! I totally feel your pain on this. Catching up on this thread now and honestly its ridiculous how much of a nightmare it is to just get decent inference without dropping a small fortune. I spent like three weeks straight tweaking configs only to have the hardware feel totally outdated the next month. It drives me crazy how these companies just keep pushing newer, heavier stuff without any regard for those of us actually trying to build local setups that dont require a professional data center budget. It feels like a scam sometimes, paying these premiums just for the privilege of keeping things offline. Anyway, before I dive into my own headaches, how big of a context window are you actually planning on hitting for your code? That usually changes the whole game for me, and I'm curious if you're aiming for full repo analysis or just single file completions.
Just wanted to say thanks for everyone chiming in. Super helpful discussion.
Yo, ngl I'm still learning the ropes with DeepSeek but I've built rigs for years. I think those R1 models are way heavier on VRAM than the 67B ones. IIRC you're gonna need massive memory for smooth inference.
Late to the party but I've been running these big quants for years. Ngl, if you want to run the heavy R1 stuff without the headache of heat and power spikes from multiple consumer cards, check out the NVIDIA RTX A6000 48GB. In my experience, single-card stability wins every time when youre deep in a coding session. That 48GB VRAM pool handles the 67B version perfectly and can squeeze in decent quants of the R1 models without breaking a sweat. If you want a quieter life and need to run the larger quants, look for a refurbished Apple Mac Studio M1 Ultra 128GB RAM. I've tried many setups over the years and for VRAM capacity per dollar, Apples unified memory is basically the easiest route for local inference right now. Just keep in mind NVIDIA is still king if you ever want to fine-tune your own models. But for just running DeepSeek while you work? The Mac is dead silent and just works.