Cloudflare has released its first machine-learning models, Clef and Clef-flash, for tasks such as choosing from predefined options. TokenPost reports that both are available through Workers AI, with weights under the Apache 2.0 licence for download and self-hosting. The practical draw is speed and flexibility, though the hosted Clef price is notably higher than the comparison model’s.
Watch Desk analysis
What happened
Cloudflare released Clef and Clef-flash on 1 October, according to TokenPost. They return probabilities for predefined choices rather than open-ended text, and their API is compatible with TypeSafe’s Jev, also known as System One. TokenPost says Clef is based on Qwen3.8-27B and Clef-flash on Qwen3.5-9B.
In the ten quality benchmarks cited by TokenPost, one of the Cloudflare models led seven, Jev led two and a model based on Google’s DiffusionGemma led the remaining test. On BANKING77, a customer-service intent classification benchmark, Clef scored 94.20 against Jev’s 79.74. Jev scored higher on When2Call and BRIGHT. Those are benchmark results reported by TokenPost, not a guarantee of performance on your own workload.
Key findings
- Fast responses
TokenPost reports median latency of 209.3 milliseconds for Clef and 38.8 milliseconds for Clef-flash. - A substantial price gap
Hosted Clef input costs $0.24 per million tokens, against $0.042 for Jev; Clef-flash costs $0.09. - Weights available to download
Both models are offered under Apache 2.0, allowing users to self-host rather than rely only on Cloudflare’s hosted service. - Results vary by task
Clef led Jev on BANKING77, but Jev came out ahead on When2Call and BRIGHT. - More fine-tuning support planned
Cloudflare is offering hands-on help with reinforcement-learning fine-tuning, with a self-service option planned for later.
Why it matters
This is a new option for teams building classification or routing systems, where a quick choice between known categories may matter more than fluent prose. The reported latency figures, particularly for Clef-flash, make the pair worth testing for that kind of work. The benchmark spread also argues for trying the relevant tasks rather than picking a model by its best score.
Price complicates the pitch. Clef’s input rate is about 5.7 times Jev’s, according to the figures reported by TokenPost. Lower latency may be worth paying for in some applications; in others, it is an expensive way to save a fraction of a second. Self-hosting changes the options, but it does not make operating a model free.
Our read
Cloudflare has put a credible new set of tools on the table, and the combination of downloadable weights, a compatible API and low reported latency is worth attention. Treat the benchmark results as a shortlist for your own tests, not a trophy cabinet. If you are evaluating them, compare end-to-end cost and accuracy on your actual workload, not just the speed of the model response.
What to watch
- Whether Cloudflare publishes more detail about the benchmark methods and test conditions.
- How the models perform on customer workloads beyond the reported benchmarks.
- When self-service reinforcement-learning fine-tuning becomes available.
- Whether users find the hosted speed worth the price premium over Jev.
Discussion spark: For a classification task, would you pay more for Clef’s reported speed, or choose Jev’s lower input price unless your own tests show a meaningful difference?
Sources and evidence
- Cloudflare’s Clef Model Posts Less Than Half Jev’s Median Latency – tokenpost.com (6 October 2026, 08:03 UTC)
Watch Desk is operated by WittyWires as an independent cross-cutting AI news tracker. It does not speak for the organisations or people it covers.