Halogen 0.17.2 Boosts Qwen Flash Next Inference to 45 TPS on AMD Strix Halo
The Halogen inference engine released version 0.17.2, enabling the Qwen 3.8 Flash Next model (~177B parameters) to sustain decode speeds of around 45 tokens per second even at high context lengths on a 128GB AMD Strix Halo system. Achieving fast local inference on massive models of nearly 177 billion parameters makes complex AI workflows more practical on workstation APUs without relying on expensive enterprise cloud GPUs. It highlights rapid optimizations in local LLM engines tailored for modern high-bandwidth unified memory architectures. The speed test was performed on a 128GB AMD Strix Halo configuration running Halogen 0.17.2 with Qwen 3.8 Flash Next. The user reported that decode throughput remained stable near 45 tokens per second during high-context usage and delivered high-quality planning outputs.
## BACKGROUND
AMD Strix Halo (Ryzen AI MAX) APUs feature unified memory architectures that offer high memory bandwidth, making them well-suited for running large AI models locally. Inference engines like Halogen continuously update execution kernels and memory access paths to increase generation throughput, measured in tokens per second (TPS).