Qwen3.8-Flash-Next (125B) Runs at 59 tok/s on AMD Strix Halo Mini PC
Yamz Labs released open-source 95 GB EXL3 quantized weights and their new Kyojin inference engine to run Qwen3.8-Flash-Next (125B MoE, 6B active) locally. Running on a single 128 GB AMD Strix Halo mini PC, the setup achieves decode speeds between 44 and 59 tok/s using speculative decoding. This demonstrates that massive 125-billion-parameter MoE models can execute at practical, production-ready speeds on consumer APU hardware with unified memory. It highlights rapid open-source optimization progress, bringing high-fidelity local LLM inference closer to individual developers. Without speculative decoding, the engine generates 32.7 tok/s, while prefill speeds remain flat near ~1,400 tok/s across context lengths up to 128K. The Kyojin engine prioritizes model accuracy, achieving a 94.1% top-1 agreement with the original FP8 baseline while guaranteeing that speculative decoding outputs match non-speculative results exactly.
## BACKGROUND
The AMD Strix Halo (Ryzen AI Max+ 395) is an APU featuring unified memory, allowing its integrated graphics hardware to access up to 128 GB of fast RAM without discrete GPU VRAM bottlenecks. Mixture-of-Experts (MoE) models only activate a portion of their total parameters (such as 6B out of 125B) for each token, offering high capability at reduced computation. Speculative decoding speeds up inference by having a smaller draft model propose candidate tokens that the larger target model verifies in parallel.