https://i.redd.it/041h6rq7t3oh1.gif
Been running a heterogeneous home cluster for a while — an old Acer laptop (12GB, CPU-only) as the primary API server, with a Windows box (RTX 3060 CUDA) and a Mac Mini (Metal) lending capacity over the network.
I wrote the orchestration on top of llama.cpp's `ggml-rpc` backend. It handles mDNS discovery, memory-aware sharding, and node health polling so the whole thing doesn't fall over if a node drops offline mid-generation. It exposes a standard OpenAI-compatible API (`/v1/chat/completions`).
Put the primary role on the weakest machine (the Acer) on purpose to see if it'd actually hold up orchestrating the API and offloading the heavy tensor math. It did — ran a full Qwen3.5 13B at ~12 tok/s, purely by borrowing VRAM/RAM from the CUDA and Metal nodes.
Repo: https://github.com/trademav/ramdeck-core-public
No GUI in this repo, API/CLI only, so you can actually read what's touching your network before running it. Obviously it doesn't beat a dedicated GPU rig on speed, that's not the point — it's for fitting models that you otherwise don't have the VRAM for. Included a built-in benchmark script so you can verify the numbers on your own hardware instead of trusting mine.
One heads up: it's source-available (Apache 2.0 + Commons Clause), not strictly OSI open source. It blocks commercial SaaS/resale, but personal/homelab use is fine. Didn't want that buried in a LICENSE file.
Video of the actual setup if you want to see it before reading code: https://youtu.be/tQgA2PmTc1g
Tear it apart, I'd rather hear the feedback now.
Source: r/homelab · by /u/Medicine_Blogscanner
