Cloudflare Releases Decision Models Clef and Clef-Flash for Agentic AI
Over recent weeks, decision models such as Typesafe AI's Jev System One have gained attention as a new AI paradigm. Unlike traditional classifiers, decision models produce bounded structured outputs quickly and consistently, enabling deterministic decisions in workflows without constant retraining. This contrasts with Large Language Models, which are non-deterministic but capable of open-ended reasoning and agentic workloads.
Cloudflare today announced two decision models, Clef and Clef-flash, hosted on Workers AI. Clef currently leads when evaluated against the Jev Decision Index, with results viewable on the live benchmark demo site. Both models are Jev-API compatible and faster than competing options. Cloudflare is open-sourcing both models under an Apache 2.0 license on Hugging Face, allowing developers to run them locally.

Cloudflare is also launching a reinforcement learning product that enables customers to fine-tune Clef for specific use cases.
Decision models classify inputs to help agents decide how to act, returning typed answers with probabilities. For instance, a customer support message can be classified for urgency and routed to the appropriate team with a probability score. This enables agents to make programmatic decisions and take actions without human intervention, deferring to humans only when necessary.
Cloudflare has tested Clef on its Threat Intelligence team to classify website domains. By analyzing a domain, Clef quickly identifies its categories—for example, classifying a domain as 95% likely fashion and 85% likely ecommerce.
Clef Technical Specifications
Rather than generating text, they accept text, JSON, images, or video plus typed questions, then return probabilities for the permitted answers. Both models support a 64K-token hosted context window, image classification, and up to 64 questions per request. Clef is the precision-oriented model, while Clef-Flash targets latency-sensitive decisions. Cloudflare reports median latency of 209.3 ms for Clef and 38.8 ms for Clef-Flash, with p95 latency of 238.6 ms and 122.4 ms, respectively. Clef is based on a Qwen 3.8 27B backbone, while Clef-Flash uses a Qwen 3.5 9B backbone.
| Specification | Clef | Clef-Flash |
|---|---|---|
| Model size | 27B | 9B |
| Base model | Qwen 3.8 27B | Qwen 3.5 9B |
| Model type | Multimodal decision model | Multimodal decision model |
| Primary use | Highest-precision decisions | Latency-critical decisions |
| Inputs | Text, JSON, images, video | Text, JSON, images, video |
| Context window | 64K tokens | 64K tokens |
| Questions per request | Up to 64 | Up to 64 |
| Question types | noul, choice, score | noul, choice, score |
| Vision encoder | Yes | Yes |
| Median latency | 209.3 ms | 38.8 ms |
| p95 latency | 238.6 ms | 122.4 ms |
| Output | Typed decisions and probabilities | Typed decisions and probabilities |
| API compatibility | Jev-compatible | Jev-compatible |
| Hosting | Workers AI | Workers AI |
| License | Apache 2.0 | Apache 2.0 |