← News
News

Cloudflare open-sources Clef and Clef-flash: end-to-end decision models on Qwen, emitting probabilities instead of text

Cloudflare released two Apache-2.0 decision models: Clef (27B) and Clef-flash (9B, ~39ms latency). They emit probabilities instead of free text, target AI agent orchestration and end-to-end decisions, and are compatible with TypeSafe AI's Jev API.

Source: AI 日报(ai6s.net)

What it is

On October 3, Cloudflare published two decision models on its blog: Clef (27B) and Clef-flash (9B), under Apache 2.0. Both are based on the Qwen architecture, but their training objective differs from chat models — they emit probability distributions instead of free text.

Key design

  • Clef (27B): flagship, optimized for accuracy
  • Clef-flash (9B): lightweight, ~39 ms first-token latency, targeting online orchestration and high-frequency calls
  • API compatibility: compatible with TypeSafe AI's Jev API; Cloudflare has already migrated parts of its Bot Manager classification pipeline internally

Why it matters

Traditional LLMs used for decisions require "generate text → parse back to structure", which is unfriendly to token budget and latency. Clef compresses the chain to "prompt → probability distribution → decision", which fits AI agent orchestration — echoing the early-October trend of OpenAI, Google and others putting "agent infrastructure" on the table.

Local deployment angle

Qwen-based architecture means native support in GGUF / llama.cpp. The 27B variant quantizes to Q4_K_M on a 24 GB consumer card; the 9B variant runs on a single mainstream GPU. We'll track community-quantized builds in the data section.