Cloudflare Releases Clef-omni, an Open-Weight Multimodal Decision AI Model
Cloudflare has introduced Clef-omni, an open-weight multimodal decision model capable of natively processing text, image, audio, and video inputs within a single API call. Built upon Qwen3-Omni-30B-A3B-Instruct, the model's weights are freely available on Hugging Face alongside hosted deployment on Cloudflare Workers AI. By natively accepting audio and video formats without requiring separate speech transcription or image-extraction pipelines, Clef-omni significantly simplifies fast multimodal decision-making workflows. Its open-weight release allows developers to host or fine-tune high-performance structured decision systems locally or on edge networks at low cost. Clef-omni operates on a 30B-parameter Mixture-of-Experts (MoE) backbone with 3B active parameters, delivering median latencies of ~130 ms for text, ~150 ms for images, and ~1.5 s for 21-second audio-video files. Designed specifically for structured decision outputs rather than standard text generation, it is priced at $0.15 per million input tokens on Cloudflare Workers AI.
## BACKGROUND
Decision AI models differ from standard conversational language models by focusing on structured task execution and schema-based evaluation rather than open-ended dialogue. In AI, open-weight models provide downloadable model parameters for self-hosting while keeping training details proprietary, and native multimodal architectures process raw audio and video without converting them to static text or image sequences first.