llama.cpp Release b11371 Adds Text-Only Support for Cloudflare's Clef Model
llama.cpp release b11371 introduces text-only support for Cloudflare's Clef decision model via pull request #29831. The update also includes static graph performance enhancements and updates to the gguf-py package. Developers running local AI pipelines can now execute Cloudflare's Clef decision models directly through llama.cpp's hardware-accelerated C/C++ inference engine. This expands open-source execution options for agentic workflows and structured decision-making without relying on cloud APIs. This release adds support specifically for the text-only capabilities of Clef, meaning multimodal inputs such as images are not yet supported. Binaries are provided across platforms including macOS, Linux, Windows, Android, and Snapdragon devices using backends such as CUDA, Vulkan, and OpenVINO.
## BACKGROUND
llama.cpp is a widely used C/C++ open-source inference engine optimized for running Large Language Models efficiently on consumer hardware. Clef is an open-weight decision model developed by Cloudflare and based on Qwen, designed to convert system states and typed schemas into structured decision outputs.