llama.cpp Release b11540 Adds SYCL Acceleration for MXFP4 MoE Models
llama.cpp release b11540 introduces hardware acceleration for MXFP4 Mixture-of-Experts (MoE) models on the SYCL backend via PR #29809. The update improves performance by utilizing arithmetic decoding and weight reordering techniques. This improvement enables faster inference for low-precision 4-bit MoE models on heterogeneous hardware supported by SYCL, including Intel GPUs. As microscaling formats like MXFP4 become standard for large AI models, efficient backend kernels are essential for low-latency execution. The optimization specifically targets MXFP4-quantized MoE layers within the SYCL compute path. Pre-built binary releases for b11540 cover platforms across macOS, Linux, Windows, Android, and Snapdragon devices.
## BACKGROUND
SYCL is an open, cross-platform C++ programming standard maintained by the Khronos Group for heterogeneous parallel computing across GPUs, CPUs, and accelerators. MXFP4 (Microscaling FP4) is a 4-bit floating-point data format designed to significantly reduce memory bandwidth and storage requirements for large language models while maintaining accuracy.