~/LLM INFERENC/developer-builds-2-800-256gb-rig-using-radeon-pro-v620-gpus-and

Developer Builds $2,800 256GB Rig Using Radeon Pro V620 GPUs and Custom vLLM Fork

A developer built a budget 256 GB VRAM inference rig using eight enterprise AMD Radeon Pro V620 cards for around $2,800. By leveraging AI to write custom RDNA2 kernels for a vLLM fork, they achieved over 3,000 tokens/second prefill and 60–100 tokens/second decode speeds on large language models. High-VRAM setups required to run large open-source LLMs locally are typically prohibitively expensive due to high enterprise GPU prices. This project demonstrates that legacy cloud gaming hardware can achieve top-tier inference performance when optimized with custom software pipelines. The system runs vLLM with Pipeline Parallelism (PP=4) without Tensor Parallelism, applying W4A16 quantization for routed experts while keeping other layers at BF16. It achieved an 800% speedup over llama.cpp, which previously struggled with poor prefill speeds (350–450 t/s) and lack of concurrency support on this hardware.

## BACKGROUND

LLM inference relies on two distinct steps: prefill, which processes the full prompt in parallel to construct the KV cache, and decode, which generates output tokens one by one. Advanced inference engines like vLLM optimize memory and throughput via PagedAttention and continuous batching, but often lack official kernel support for older enterprise AMD architectures like RDNA2.

## REFERENCES

## KEYWORDS

#LLM Inference#Hardware#vLLM#AMD ROCm#Local AI

$ subscribe --daily

Developer Builds $2,800 256GB Rig Using Radeon Pro V620 GPUs and Custom vLLM Fork | Daily News