Cloudflare Releases Clef Decision Models and RL Fine-Tuning Platform
Cloudflare has launched Clef and Clef-flash, a suite of open-weights decision models available on Workers AI, alongside a new reinforcement learning (RL) fine-tuning platform. The flagship Clef model currently leads the Jev Decision Index benchmark evaluation. This marks a major move by cloud infrastructure provider Cloudflare into specialized decision AI, offering developer-friendly structured outputs instead of traditional unstructured text. It enables developers to fine-tune specialized decision models using RL directly within Cloudflare's edge ecosystem. Clef is priced at $0.24 per million input tokens, while the lightweight Clef-flash costs $0.09 per million tokens. Although marketed with open weights derived from Qwen base models, Cloudflare has not released the full training data or pipelines required to reproduce the models from scratch.
## BACKGROUND
Unlike general Large Language Models (LLMs) that generate unstructured text, decision models are optimized to rapidly produce structured, typed outputs for classification or automated choices. In machine learning, releasing open weights allows developers to run and fine-tune model parameters locally, though it differs from full open-source AI where training datasets and code pipelines are also shared. Reinforcement Learning (RL) fine-tuning helps models optimize their decision-making process based on feedback and reward signals.