Qwen 27B model can now run on AMD NPUs using FastFlowLM
Developers have demonstrated running the 27-billion parameter Qwen LLM on AMD NPUs using the FastFlowLM inference framework. However, token generation performance is currently extremely slow, reaching only around 1 token per second during decoding. This milestone expands local AI inference capabilities by bringing large 27B parameter models to consumer AMD NPU hardware for the first time. However, the sluggish 1 token per second rate highlights that consumer NPU architectures still face severe memory bandwidth bottlenecks with large LLMs. While FastFlowLM is optimized specifically for AMD Ryzen AI NPUs, running models as large as 27B pushes the hardware far beyond its sweet spot. The current throughput of 1 token per second makes this deployment useful primarily as a technical benchmark rather than a practical tool for daily conversation.
## BACKGROUND
Neural Processing Units (NPUs), such as AMD's XDNA series integrated into Ryzen AI processors, are specialized chips designed for efficient, low-power AI acceleration. FastFlowLM is an inference runtime tailored specifically to optimize local LLM execution on AMD NPUs. Large models like Qwen 27B typically demand high memory bandwidth and capacity, which consumer NPUs often struggle to deliver compared to high-end dedicated GPUs.