llama.cpp b11294 Release Adds Row Prefetching for Lazy Model Loading
The llama.cpp project released version b11294, introducing row prefetching capabilities (llama_prefetch_rows) for lazy model loading, including cross-platform support for Windows. It also includes code structure cleanups and clarifies padding token definitions for Gemma 4 models. Row prefetching helps optimize memory access patterns during lazy model loading, reducing latency and I/O bottlenecks when executing Large Language Models. Extending this functionality to Windows ensures consistent performance improvements across major operating systems. Row prefetching is strictly enabled when running in lazy mode and routes through llama-impl rather than exposing llama-mmap directly to model logic. The Windows implementation was contributed by @praneshgo and merged alongside build fixes and Gemma 4 token adjustments.
## BACKGROUND
llama.cpp is a popular open-source C/C++ library designed for efficient inference of Large Language Models on local hardware. Lazy model loading defers reading model weights into memory until needed, while prefetching proactively reads upcoming data to prevent performance stalls.