~/FPGA/running-qwen3-5-llm-inference-on-cheap-repurposed-fpga-mining-hardware

Running Qwen3.5 LLM Inference on Cheap Repurposed FPGA Mining Hardware

Developer Nero7991 successfully implemented custom RTL inference for Qwen3.5 9B and 27B INT4 models in VHDL on repurposed SQRL FK33 and Jungle Cat FPGA mining boards. Running on recycled Xilinx Virtex UltraScale+ hardware equipped with HBM2 memory, the system achieved generation speeds up to 3.2 tokens/second at a 75 MHz clock rate. This project proves that low-cost, repurposed crypto-mining boards with High Bandwidth Memory (HBM2) can serve as viable hardware accelerators for LLM inference. It highlights an accessible open-source path toward bypassing expensive commercial GPUs by implementing custom logic directly on secondhand enterprise FPGAs. The open-source repository (llm.vhdl) handles quantized INT4 inference across multi-board setups using pipeline parallelism and host-managed residual pass-through. Projections show that scaling the design to a 200 MHz clock rate across four XCVU35P dies (32GB HBM2 total) could achieve generation speeds of 25 tokens/second for short contexts and support context lengths up to 262k tokens.

## BACKGROUND

FPGAs (Field-Programmable Gate Arrays) are reconfigurable semiconductor devices that allow developers to design customized hardware circuits using hardware description languages like VHDL. Mining boards such as the SQRL FK33 and Jungle Cat feature Xilinx Virtex UltraScale+ FPGAs paired with High Bandwidth Memory (HBM2), providing up to 400 GB/s of memory bandwidth which is essential for memory-bound LLM inference tasks.

## REFERENCES

## KEYWORDS

#FPGA#Hardware Acceleration#LLM Inference#Qwen#Open Source AI

$ subscribe --daily