Clef-omni handles audio and video without needing transcription pipelines
Cloudflare just expanded its open-weight decision model lineup with Clef-omni, which natively processes audio (wav, mp3) and video (mp4, webm) alongside standard text and images. Unlike traditional LLMs, these decision models aren't designed to generate conversational text; instead, they provide schema-constrained scoring. By integrating multimodal inputs directly, the model removes the need for cascading pipelines where you would normally have to transcribe speech-to-text or separate audio tracks from video frames before making a decision.
How Clef-omni processes multimodal data
The architecture is built on a Qwen3-Omni-30B-A3B-Instruct mixture-of-experts (MoE) foundation. While the base model can handle various modalities, Cloudflare stripped away the text-to-speech output components to focus purely on decision-making. When a payload is sent to the model, it performs a rapid prefill pass, scoring all valid parameter options and modalities at once.
Because it skips the token generation phase entirely, there is no overhead from captioning or transcribing files. Media elements are mapped into a unified sequence where audio and video remain synced with visual frames. The model then uses a two-stage attention routing system to pull candidate values from internal embeddings. This means the model gathers evidence from text, visuals, or audio streams simultaneously, and then uses field vectors to cross-attend across the full context to calculate confidence scores. To ensure the results stay within a specific format, a built-in lexical grammar preserves the semantics of the options, allowing for fast, schema-bound scoring.
Training and technical specifications
The training process for Clef-omni mirrored the original Clef release. The developers froze the Qwen3 backbone and utilized low-rank adapters (LoRA) for training. To ensure the model remains resilient against changes in prompt structure, field ordering, or schema variations, they applied a post-training approach that combines Brier score calibration with label-smoothed cross-entropy loss.
Beyond the omni version, there are two other key updates to the family:
- Clef: Now operates at a faster speed than previous iterations.
- Clef-flash: The price has been reduced, making it more affordable than the Jev model from TypeSafe.
Integrating Clef-omni into a workflow
If you are moving from a text-only decision model to Clef-omni, the primary shift is in how you handle your input pipeline. You no longer need a separate Whisper or similar speech-to-text model to "preprocess" your audio before asking the model to categorize or score it. You can feed the raw media file directly.
To get started with the multimodal inputs, you can find the open weights on HuggingFace or refer to the developer documentation. The lack of output token generation is the critical detail here; if your project requires a chatty response, this isn't the right tool. But if you need a high-confidence score on a specific parameter based on a video clip and a text prompt, this architecture is significantly more efficient than a standard multimodal LLM.
The speed of iteration on this project is also a point of interest. The original Clef models were developed in less than a week—conceived on a Friday evening, trained over a weekend, and launched by the following Thursday. This suggests that the decision-model framework is highly modular, allowing for rapid adaptation of different foundations (like Qwen3) into specialized scoring tools.

Does the native video processing handle long clips without aggressive chunking? I'd love to know the max duration supported for accurate schema scoring.