[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"$f3nopmse7pnfdc":3},{"slug":4,"category":5,"publishedAt":6,"titleZh":7,"titleEn":8,"summaryZh":9,"summaryEn":10,"models":11,"hardwares":12,"bodyZh":13,"bodyEn":14,"sourceName":15,"sourceUrl":16,"updatedAt":17},"cloudflare-clef-decision-models","NEWS","2026-10-04T05:00:00.000Z","Cloudflare 开源 Clef 与 Clef-flash：Qwen 架构的端到端决策模型，输出概率不输出文本","Cloudflare open-sources Clef and Clef-flash: end-to-end decision models on Qwen, emitting probabilities instead of text","Cloudflare 发布两款 Apache 2.0 开源决策模型：Clef（27B）与 Clef-flash（9B，~39ms 延迟），输出概率而非自由文本，专为 AI 智能体编排与端到端决策设计，与 TypeSafe AI 的 Jev API 兼容。","Cloudflare released two Apache-2.0 decision models: Clef (27B) and Clef-flash (9B, ~39ms latency). They emit probabilities instead of free text, target AI agent orchestration and end-to-end decisions, and are compatible with TypeSafe AI's Jev API.",[],[],"### 是什么\n\n10 月 3 日，Cloudflare 在官方博客发布两款决策模型 `Clef`（27B）与 `Clef-flash`（9B），采用 **Apache 2.0 协议开源**。两款模型都基于 Qwen 架构改造，但训练目标与传统对话模型不同——**输出的是概率分布，而非自由文本**。\n\n### 关键设计\n\n- **Clef（27B）**：旗舰版本，主打精度\n- **Clef-flash（9B）**：轻量版本，首 token 延迟约 39 毫秒，面向在线编排与高频调用\n- **API 兼容**：与 TypeSafe AI 的 Jev API 兼容，已在 Cloudflare 内部替代部分 Bot Manager 分类链路\n\n### 为什么值得关注\n\n传统 LLM 用于决策时需要\"先编出文本再解析回结构\"，对 token 预算和延迟都不友好。Clef 把这条链路压成\"prompt → 概率分布 → 决策\"，对 AI 智能体编排场景更友好——也呼应了 10 月初 OpenAI、Google 等厂商把\"智能体基础设施\"提到台面的趋势。\n\n### 本地部署角度\n\n基于 Qwen 架构意味着 GGUF\u002Fllama.cpp 一线工具链原生支持；27B 版本在 24GB 显存显卡上可量化到 Q4_K_M 运行，9B 版本更是单卡消费级即可。本站后续会跟踪社区量化版的发布。","### What it is\n\nOn October 3, Cloudflare published two decision models on its blog: `Clef` (27B) and `Clef-flash` (9B), under **Apache 2.0**. Both are based on the Qwen architecture, but their training objective differs from chat models — they emit **probability distributions instead of free text**.\n\n### Key design\n\n- **Clef (27B)**: flagship, optimized for accuracy\n- **Clef-flash (9B)**: lightweight, ~39 ms first-token latency, targeting online orchestration and high-frequency calls\n- **API compatibility**: compatible with TypeSafe AI's Jev API; Cloudflare has already migrated parts of its Bot Manager classification pipeline internally\n\n### Why it matters\n\nTraditional LLMs used for decisions require \"generate text → parse back to structure\", which is unfriendly to token budget and latency. Clef compresses the chain to \"prompt → probability distribution → decision\", which fits AI agent orchestration — echoing the early-October trend of OpenAI, Google and others putting \"agent infrastructure\" on the table.\n\n### Local deployment angle\n\nQwen-based architecture means native support in GGUF \u002F llama.cpp. The 27B variant quantizes to Q4_K_M on a 24 GB consumer card; the 9B variant runs on a single mainstream GPU. We'll track community-quantized builds in the data section.","AI 日报（ai6s.net）","https:\u002F\u002Fai6s.net\u002F6ac0377c05257b0857147348.html","2026-10-04T01:02:09.609Z"]