Cloudflare 发布决策模型 Clef 和 Clef-flash,称智能体决策无需人类介入
Cloudflare says its new Clef model means humans no longer need to be in the loop for AI agents
Cloudflare 发布面向 AI 智能体的决策模型 Clef 和 Clef-flash,输出带概率的分类结果,称智能体可据此自主决策并在必要时交给人类。据其自报数据,Clef-flash 中位延迟约 39 毫秒、Clef 约 209 毫秒,均快于竞品 Jev 的 524 毫秒以上。
Cloudflare is releasing Clef and Clef-flash, two decision models for AI agents. Speed is the company's main selling point against frontrunner Jev.
A decision model returns a brief classification with probabilities rather than a long text response. Given a customer support message, Clef assesses its urgency and identifies the team that should handle it. Downstream code can use those results to route a ticket, trigger an escalation, or hand the case to a human.
According to Cloudflare, "a human does not necessarily need to be in the loop for agentic decisions anymore." Agents can "programmatically gather context, make decisions, and take actions on tasks, or defer to a human when needed." Cloudflare takes the name from music, where a clef assigns pitches to the lines of a staff. Similarly, a decision model sets the framework for the actions that follow. The phonetic resemblance to "Jev" is probably no accident either.

These "decision models" fill a niche between large language models and traditional classifiers. Language models can reason and call tools, but their outputs vary and they can be slow. Traditional classifiers are fast but need retraining for every new category. Cloudflare sees Clef as a direct answer to TypeSafe AI's Jev model and keeps the API fully compatible so customers can switch easily.
Clef-flash returns a decision in 39 milliseconds
Across 43 benchmarks, Clef and the smaller Clef-flash are faster than all relevant competing decision models, according to Cloudflare. The company reports median latency of about 39 milliseconds for Clef-flash and about 209 milliseconds for Clef, compared with just over 524 milliseconds for Jev. Both models run directly on Cloudflare's infrastructure, which the company says also lets them benefit from proximity to edge data centers.

Cloudflare's threat intelligence team is already testing Clef to classify websites. In one example, it assigns a domain a 95 percent probability of being a fashion website and 85 percent of being an online store. The probability of it being a phishing site is under one percent. Fetching, rendering, and classifying the site took 2.2 seconds. The company's fastest general-purpose language model took 4.7 seconds for the same process and returned only two categories.
Clef can also process images, according to Cloudflare, while Jev is limited to text so far. Its 64,000-token context window holds twice as much input as Jev's. Clef also leads on the Jev Decision Index in Cloudflare's own benchmarks.
| Benchmark · accuracy | Clef | Clef-flash | Jev | DiffusionGemma Jev | Kev 9B | Laya |
|---|---|---|---|---|---|---|
| API Bank | 91.93 | 93.11 | 88.19 | 83.66 | 56.30 | 11.41 |
| When2Call | 72.37 | 65.58 | 80.97 | 75.44 | 49.62 | 11.94 |
| PhishNChips | 79.60 | 75.05 | 62.55 | 85.35 | 50.75 | 50.15 |
Cloudflare is based on Qwen models
Clef is based on Qwen3.8-27B and Clef-flash on the smaller Qwen3.5-9B, according to Cloudflare. Cloudflare leaves the base models unchanged during training and uses its own synthetic data to train extra components. These components derive answer options and probabilities from the models' internal computations.
Cloudflare also uses its own variant of Reinforcement Learning for Calibrated Decisions (RLCD), the training method TypeSafe used to train Jev. RLCD trains models to answer multiple questions about an input in a single call, aiming to assign probabilities that match how often the answers are actually correct. The company previously experimented with DiffusionGemma to derive fixed decision values from a language model's internal computations.
Customers will be able to fine-tune Clef for their own tasks
Alongside the launch, Cloudflare is rolling out a reinforcement learning service that lets customers tailor Clef to their own tasks. A team of forward deployed engineers will initially handle fine-tuning with customers, with a self-service platform planned for later.
Customers will be able to build a dataset by logging requests through AI Gateway, then evaluate those requests in containers that serve as an RL sandbox. A new trainer component will let them deploy the fine-tuned model on Workers AI. To run custom models, Cloudflare uses technology from Replicate, which it acquired in late 2025.

Cloudflare says it plans to use Clef internally to review abuse reports, sort support requests, and distinguish useful bots from harmful ones. Both models run on Cloudflare's Workers AI platform and are available on Hugging Face under the Apache-2.0 license.
The approach was popularized by TypeSafe AI, the startup founded by former OpenAI researcher Diogo Almeida, which introduced Jev in mid-September. TypeSafe markets Jev as a model "without hallucinations," though that only guarantees it stays within predefined answer options and doesn't prevent it from choosing incorrectly. In late September, OpenAI followed with a Decisions API built on GPT-6 Luna that also accepts context as text or images.
Cloudflare primarily runs a global network for content delivery, DNS, and security services and has attracted little attention for its own AI models. Its recent AI headlines have focused on giving website owners control over access. In July, it let site operators block or allow AI bots based on their purpose.
AI News Without the Hype – Curated by Humans
Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section.
来源:The Decoder:AI News · the-decoder.com