Skip to main content
Cloudflare blog header with text introducing Clef decision models and a colorful crystal ball icon.

Cloudflare Releases Decision Models Clef and Clef-Flash for Agentic AI

Over recent weeks, decision models such as Typesafe AI's Jev System One have gained attention as a new AI paradigm. Unlike traditional classifiers, decision models produce bounded structured outputs quickly and consistently, enabling deterministic decisions in workflows without constant retraining. This contrasts with Large Language Models, which are non-deterministic but capable of open-ended reasoning and agentic workloads.

Cloudflare today announced two decision models, Clef and Clef-flash, hosted on Workers AI. Clef currently leads when evaluated against the Jev Decision Index, with results viewable on the live benchmark demo site. Both models are Jev-API compatible and faster than competing options. Cloudflare is open-sourcing both models under an Apache 2.0 license on Hugging Face, allowing developers to run them locally.

Cloudflare is also launching a reinforcement learning product that enables customers to fine-tune Clef for specific use cases.

Decision models classify inputs to help agents decide how to act, returning typed answers with probabilities. For instance, a customer support message can be classified for urgency and routed to the appropriate team with a probability score. This enables agents to make programmatic decisions and take actions without human intervention, deferring to humans only when necessary.

Cloudflare has tested Clef on its Threat Intelligence team to classify website domains. By analyzing a domain, Clef quickly identifies its categories—for example, classifying a domain as 95% likely fashion and 85% likely ecommerce.

Clef Technical Specifications

Rather than generating text, they accept text, JSON, images, or video plus typed questions, then return probabilities for the permitted answers. Both models support a 64K-token hosted context window, image classification, and up to 64 questions per request. Clef is the precision-oriented model, while Clef-Flash targets latency-sensitive decisions. Cloudflare reports median latency of 209.3 ms for Clef and 38.8 ms for Clef-Flash, with p95 latency of 238.6 ms and 122.4 ms, respectively. Clef is based on a Qwen 3.8 27B backbone, while Clef-Flash uses a Qwen 3.5 9B backbone.

Specification Clef Clef-Flash
Model size 27B 9B
Base model Qwen 3.8 27B Qwen 3.5 9B
Model type Multimodal decision model Multimodal decision model
Primary use Highest-precision decisions Latency-critical decisions
Inputs Text, JSON, images, video Text, JSON, images, video
Context window 64K tokens 64K tokens
Questions per request Up to 64 Up to 64
Question types noul, choice, score noul, choice, score
Vision encoder Yes Yes
Median latency 209.3 ms 38.8 ms
p95 latency 238.6 ms 122.4 ms
Output Typed decisions and probabilities Typed decisions and probabilities
API compatibility Jev-compatible Jev-compatible
Hosting Workers AI Workers AI
License Apache 2.0 Apache 2.0

cloudflare


Felipe Santos

"Artificial intelligence can process the world in milliseconds, but only the human heart can give meaning to every second lived" - Mr. Santos