~/HARDWARE/diy-cluster-of-repurposed-crypto-mining-boards-runs-distributed-llm-inference

DIY Cluster of Repurposed Crypto Mining Boards Runs Distributed LLM Inference

A Reddit user demonstrated a DIY hardware setup combining six repurposed AMD BC-250 crypto-mining boards to run distributed local LLM inference via llama.cpp with RPC and Vulkan backends over 1Gb Ethernet. The setup successfully runs large context models like Qwen Next Flash IQ2_XS at up to 100k context across four boards, while two boards run a 35B Q4 model at 60 tokens per second. This project highlights the practical viability of repurposing low-cost, e-waste mining hardware into high-capacity nodes for local AI inference. It demonstrates how open-source tools like llama.cpp RPC enable enthusiasts to aggregate memory bandwidth across budget hardware without buying expensive enterprise GPUs. The hardware setup houses five BC-250 boards in an ASRock 4U12G chassis using custom cardboard intake fans, communicating over basic 1Gb Ethernet. Benchmark metrics include 28 tokens/sec generation on Qwen Next Flash IQ2_XS (115 prompt processing tokens/sec at 50k context) and 450 prompt processing tokens/sec on 35B Q4 quants.

## BACKGROUND

The AMD BC-250 is a specialized cryptocurrency mining board featuring cut-down PlayStation 5 APUs paired with 16GB of fast GDDR6 memory. llama.cpp is a popular open-source LLM inference framework that includes an RPC (Remote Procedure Call) mechanism, allowing users to offload model layers and split compute across multiple networked nodes.

## REFERENCES

## KEYWORDS

#hardware#local-ai#llama-cpp#distributed-inference#homelab

$ subscribe --daily

DIY Cluster of Repurposed Crypto Mining Boards Runs Distributed LLM Inference | Daily News