llama.cpp Release b11493 Adds Grouped MoE XMX GEMM for Intel SYCL
Build b11493 of llama.cpp introduces grouped MoE (Mixture of Experts) XMX GEMM kernel optimizations for Intel's SYCL backend. This update comes alongside updated pre-compiled binaries spanning multiple operating systems and architectures including Linux, Windows, macOS, and Android. Grouped GEMM kernels significantly accelerate Mixture-of-Experts LLM inference by processing matrix multiplications across multiple experts in a single kernel execution. This improvement helps users running Intel GPUs via SYCL achieve better throughput when running MoE architectures like Mixtral or DeepSeek. The pull request (#29245) leverages Intel's Xe Matrix Extensions (XMX) hardware engines via the SYCL backend to execute parallel matrix math for dynamic token routing in MoE models. Pre-compiled binaries were refreshed across CUDA 12/13, ROCm, Vulkan, OpenVINO, SYCL FP16/FP32, and Snapdragon builds.
## BACKGROUND
SYCL is an open-standard, C++ based programming model developed by Khronos Group that enables heterogeneous computing across diverse hardware accelerators like Intel GPUs. Intel Xe Matrix Extensions (XMX) are dedicated hardware matrix compute engines integrated into Intel Arc and Data Center GPUs to accelerate AI workloads. Mixture-of-Experts (MoE) models route tokens dynamically to different sub-networks (experts), requiring grouped GEMM routines to process multiple distinct matrix operations efficiently in parallel.