
Cloudflare has introduced Clef, a new family of open-source decision models it says can make structured choices faster and more consistently than general-purpose large language models. The company is also launching a new reinforcement learning fine-tuning service, initially offered with hands-on support and later as a self-serve platform for customers.
Cloudflare launches Clef and Clef-flash
The company said it is releasing two Cloudflare-trained decision models, Clef and Clef-flash, hosted on Workers AI and available under an Apache 2.0 license on Hugging Face. Cloudflare describes the models as “fully Jev-API compatible,” making them straightforward to test in existing workflows designed for decision-model use cases.
Cloudflare says Clef currently leads the Jev Decision Index benchmark, and it points readers to a live demo site for full results. The larger Clef model is positioned as the more capable precision option, while Clef-flash is aimed at latency-sensitive decisions.
What Cloudflare means by a decision model
In Cloudflare’s framing, a decision model is not meant to generate open-ended text the way an LLM does. Instead, it classifies inputs and returns bounded, structured outputs with probabilities that software can act on directly. That can be useful for routing support tickets, escalating urgent issues, or sending tasks to a human when confidence is low.
The company argues that this approach is especially well suited to agentic workflows, where systems need to gather context, make a decision, and then execute the next step programmatically. Rather than asking a model to write an explanation first, a decision model can produce the classification needed by the application immediately.
Why Cloudflare built Clef
Cloudflare said it has been testing Clef internally on its Threat Intelligence team to classify website domains. In one example, the company said Clef, using Browser Run, could identify likely categories for a domain — such as fashion, ecommerce, or phishing — in 2.2 seconds to fetch, render, and classify the site.
For comparison, Cloudflare said its fastest general LLM, gpt-oss-120b, took 4.7 seconds in the same workflow and returned only two classifications. Cloudflare presented the latency and output differences as an example of why compact decision models can be useful in threat intelligence and other programmatic decision-making systems.
The name “Clef” is a reference to music theory, where a clef sets the pitch reference for a staff. Cloudflare said it chose the name because the model helps define the context for the actions that follow, and the “CF” also nods to Cloudflare.
How Clef differs from other decision models
Cloudflare highlighted several differentiators for Clef. It says the model includes a vision encoder, allowing it to classify images as well as text, unlike Jev, which Cloudflare says currently supports text classification only.
Clef also has a 64k context window, which Cloudflare says is twice the size of Jev’s 32k window. That, the company said, lets users provide more state for the model to evaluate before it makes a decision.
On performance, Cloudflare said Clef scores competitively across a range of benchmarks tied to decision-making tasks. It published results for a mix of evaluation sets, including:
- BFCL · case exact — Clef: 98.47, Clef-flash: 98.76
- API-Bank · accuracy — Clef: 91.93, Clef-flash: 93.11
- BANKING77 · macro-F1 — Clef: 94.20, Clef-flash: 90.93
- CLINC150+OOS · macro-F1 — Clef: 97.43, Clef-flash: 66.77
- PhishNChips · accuracy — Clef: 79.60, Clef-flash: 75.05
Cloudflare also said it compared Clef against Typesafe’s eval suite and that its models beat Jev in three of four areas. The company noted that Clef-flash performed especially well relative to its speed.
Latency is a major part of the pitch
Cloudflare says the models are designed to be fast not just because of their architecture, but also because they run on Workers AI close to the edge. That infrastructure, the company said, can reduce network latency and make the models suitable for hot-path decisions in agent workflows.
In a broader benchmark comparison across 43 evaluations, Cloudflare said its Clef models posted lower latency than other decision models, with one exception it called out as very fast but more quality-limited. The published figures were:
- Median latency — Clef: 209.3 ms, Clef-flash: 38.8 ms
- p95 latency — Clef: 238.6 ms, Clef-flash: 122.4 ms
The company says the models return strictly typed outputs and are fully API-compatible, which should make adoption easier for teams already using decision-model-style interfaces. Cloudflare also said customers can trust that it does not read, store, or train on requests or responses unless they choose to use the fine-tuning product.
How Cloudflare trained Clef
Cloudflare said Clef builds on earlier experiments where it adapted a DiffusionGemma model to output deterministic probabilities. For Clef, however, the company said it uses Qwen as the base model and post-trains it for decision-model tasks.
During inference, Cloudflare says Clef performs a prefill-only pass and then scores valid schema choices in parallel, rather than generating text token by token. The company says that non-autoregressive setup is a key reason the model can be faster than typical LLMs.
Cloudflare said it froze Qwen3.8-27B for Clef and Qwen3.5-9B for Clef-flash, then optimized a routing head alongside rank-256 low-rank adapters. It also said training used label-smoothed cross-entropy, Brier loss for probability calibration, and internal synthetic datasets that permuted field order, prompts, and schema structures.
The company said it also developed Reinforcement Learning for Calibrated Decisions, or RLCD, as a secondary optimization target. Cloudflare described that as giving partial credit to adjacent ordinal choices, rewarding precise outputs, and applying a reference penalty to limit distribution shift.
Fine-tuning is the next step
Alongside the model release, Cloudflare is introducing a reinforcement learning service for custom fine-tuning. The first phase will involve customers working with Cloudflare’s forward-deployed engineer team, with a self-serve platform planned later.
Cloudflare said the goal is to support highly specific internal and customer workflows, including Trust & Safety submission review, support ticket triage, and bot classification. The company said its long history of network data gives it a base for building specialized models tuned to those domains.
Its planned RL workflow uses several parts of Cloudflare’s AI stack:
- AI Gateway to capture AI traffic and build datasets from requests
- Workers AI to generate rollouts against the base Clef model
- Containers as an RL sandbox for scoring and replaying agent actions
- Trainer to update weights for fine-tuned Clef models
- Workers AI + BYO Model to redeploy the fine-tuned model
Cloudflare said this is part of its broader AI platform work and ties into infrastructure it has been building for captured traffic, sandboxed experimentation, and bringing custom models onto Workers AI.
What Cloudflare is signaling
The launch suggests Cloudflare wants to position itself not only as an inference provider, but as a place to build and deploy specialized decision systems for agentic workloads. That includes a public open-source model release, benchmark claims against competitors, and a managed path toward RL-based fine-tuning.
For now, Cloudflare says Clef is available on Workers AI and on Hugging Face, and it is inviting customers with relevant use cases to engage as design partners. The company also says its machine learning team is hiring.
Source: Original report
Was this helpful?
Explore more: AI Automation Services More AI & Automation Tech News
Last Modified: October 1, 2026 at 10:35 pm
0 views
