Scaling a Home Lab to Run Frontier-Class Open AI Models
A local AI enthusiast detailed their journey scaling home AI infrastructure from a single RTX 3090 to a 20-node DGX Spark cluster across neighboring houses. This setup enables local, offline inference for massive open-weights models like Kimi K3 (2.8T) and DeepSeek 671B. The project demonstrates that individual developers can run frontier-class open models completely offline for data privacy and long-term cost control. It also highlights real-world physical boundaries of extreme home labs, including power grid limits, thermal issues, and high-speed inter-node bandwidth. The hardware evolved from offloading to CPU DDR4 RAM to a 16x RTX 3090 rig drawing 6kW that repeatedly tripped circuit fuses, before migrating to power-efficient ASUS GB10 (DGX Spark) nodes. Connected via a 100Gbit network and running optimized vLLM and SGLang images, the combined cluster achieved up to 20 tokens per second even at a 300k context window.
## BACKGROUND
Running massive language models locally requires immense compute capacity and memory bandwidth. Architectures like Mixture of Experts (MoE) reduce computation by activating only a subset of parameters per token, enabling models like DeepSeek 671B to run more efficiently. Additionally, LLM inference relies on a compute-heavy prefill phase to process context prompts followed by a memory-heavy decode phase, making high-speed inter-GPU links crucial for processing long context windows.