More on the topic…
Andrew Chen walks through his homelab setup for running AI models locally. The centerpiece is a Framework Desktop with an AI Max+ 395 processor, paired with a 5090 eGPU that handles Qwen 3.8 27B and pushes over 150 tokens per second. He's also got two DGX Spark machines running Deepseek v4 Flash for heavier workloads, a Raspberry Pi 5 for monitoring, and a Mac mini he uses as his main development box. Everything fits into a 10-inch DeskPi rack, though he admits the eGPU purchase is something he regrets and wouldn't recommend to others.
The setup runs Hermes as his default interface with a homegrown routing plugin he built called Arch-Router. This router decides on the fly whether to handle requests locally or bump them to cloud providers and frontier models. Right now he's hitting about 60-70% local execution, with a goal to eventually get to 100% local processing. The two Spark machines handle batch jobs and background tasks—cron jobs, longer builds, that kind of thing that doesn't need instant feedback.
He's clear-eyed about the whole thing: you don't need this much hardware to experiment with local AI. He started with just the Mac mini and kept adding pieces over time because he could. The setup works for his use case, but it's not a template anyone should blindly copy. The real lesson is that you can start small and scale up as your needs become clearer.
Questions about this article
No questions yet.