llama.cpp Release b11382 Adds Float16 Support for WebGPU Operations
llama.cpp release b11382 introduces 16-bit floating-point (f16) support for the `fill` and `set_rows` tensor operations within its WebGPU backend. This update improves data precision handling and memory performance for browser-based LLM inference. Expanding float16 coverage in the WebGPU backend allows open-source language models to run more efficiently directly inside web browsers. It reduces runtime memory footprint and speeds up compute-heavy tensor operations without requiring complex native CUDA setup. The patch (PR #29897) enables half-precision float calculations for matrix filling and row setting operations on GPU buffers. Official binary builds for build b11382 were published across Linux, macOS, Windows, Android, and Snapdragon platforms.
## BACKGROUND
llama.cpp is a widely used open-source C/C++ library designed for fast inference of Large Language Models (LLMs) on consumer hardware. WebGPU is a modern Web standard providing high-performance, low-level access to graphics processing units directly inside browser environments. Using 16-bit floating-point precision (f16) instead of 32-bit float (fp32) halves memory bandwidth consumption and significantly accelerates AI workloads on modern GPUs.