Developer Tests Cloudflare's Clef Model Locally with llama.cpp and Reports Poor Results
A developer tested Cloudflare's open-source Clef decision model locally using llama.cpp and reported poor classification accuracy on basic test prompts. Running a Q4_K_M quantization of the model, the user observed significantly lower confidence scores compared to alternative models like Jev. Cloudflare recently introduced Clef as a high-speed multimodal decision model designed for structured classification and agentic workflows. These early local benchmarks suggest that local quantization or self-hosted inference setups like llama.cpp may face degradation issues compared to cloud-native implementations. When evaluating a simple billing support prompt, Clef's Q4_K_M GGUF model yielded a boolean 'noul' confidence score of only 0.57 for a refund request, whereas the Jev model scored 0.99. Running via llama.cpp 0.6.0 over 334 input tokens, the model also produced low overall confidence scores across team and urgency classification tasks.
## BACKGROUND
Cloudflare's Clef is a 27B multimodal decision model that converts structured question schemas and input states into probabilistic decisions. Quantization formats like Q4_K_M compress model weights into 4-bit representations to allow large language models to run on local consumer hardware, though this can sometimes impact precision on specialized tasks.