All News
cloudflareclefjevtypesafe-aidecision-modelsopen-weights

Cloudflare ships Clef, an open-weight answer to Jev, with a price critics noticed

Cloudflare released Clef and Clef-flash, open-weight decision models on Workers AI, with vendor-reported wins over Jev and a price that drew immediate pushback.

Vlad MakarovVlad Makarovreviewed and published
5 min read
Cloudflare ships Clef, an open-weight answer to Jev, with a price critics noticed

Cloudflare published Introducing Clef: our open-source decision models, and new RL fine-tuning platform on October 1, 2026: two open-weight "decision models", Clef and Clef-flash, served on its Workers AI platform, plus a reinforcement-learning fine-tuning product. The framing is the one TypeSafe AI introduced two weeks earlier with Jev — bounded, typed classifications instead of prose. The packaging is not. Cloudflare is shipping Apache 2.0 weights, adding a vision encoder, and putting a price on the product that Hacker News commenters immediately set against the incumbent's.

Two models behind a frozen backbone

Clef is post-trained from Qwen3.8-27B and Clef-flash from Qwen3.5-9B, and Cloudflare's own description is unusually explicit about where the work went: "By freezing Qwen3.8-27B for Clef and Qwen3.5-9B for Clef-flash, we jointly optimized the routing head alongside rank-256 low-rank adapters." At inference the Qwen backbone performs a prefill-only pass, after which the choices are scored in parallel, which is the mechanism behind the latency numbers below. Both models are published on Hugging Face under Apache 2.0 and hosted on Workers AI.

The name is a music-theory reference rather than a codename. A clef is the symbol that assigns pitch names to a staff, and Cloudflare notes that "the CF hearkens to Cloudflare."

What a decision model is, in the vendor's telling

Cloudflare defines the class narrowly on purpose: a model that makes bounded, typed classifications with probabilities attached, so that surrounding code can route, escalate, or hand a case to a human. Is this support message urgent, which team should own it. The output is a choice from a fixed set, not a sentence, and the company positions that against open-ended LLM calls, which it calls non-deterministic for the same job. TypeSafe made the identical argument for Jev, so the two vendors now differ mostly on capacity. Cloudflare points to its vision encoder, which lets Clef classify images where Jev is text-only, and to a 64,000-token context window against Jev's 32,000. It also claims Clef currently leads the Jev Decision Index, keeps a live demo open to inspection, and says both models are "fully Jev-API compatible" — a detail aimed squarely at anyone weighing the cost of switching.

The quality table is Cloudflare's arithmetic

Every figure below is self-reported by Cloudflare. No third-party run is cited anywhere in the launch material, and the Jev column is the vendor's own testing of its rival.

BenchmarkClefClef-flashJev
BFCL case exact98.4798.7695.75
ToolRet nDCG@1069.1966.4365.28
API-Bank accuracy91.9393.1188.19
Home appliances case exact82.9597.7352.27
When2Call accuracy72.3765.5880.97
BANKING77 macro-F194.2090.9379.74

Read the rows in both directions. Clef-flash wins the tool-retrieval and API-bank columns, and both Cloudflare models crush a home-appliance classification set where Jev scores 52.27. But on When2Call the ordering inverts: Jev's 80.97 outruns both Cloudflare models, and Clef-flash trails at 65.58. A vendor table that contains a rival win is not automatically honest, but it is easier to audit, and it is the row a skeptical reader should start with. Cloudflare also reports beating Jev in three of four areas on TypeSafe's own WorkflowEvals, including invoice processing at 64.7 against 61.8 and security incidents at 62.9 against 61.7, while conceding Jev wins agent-trace observability at 71.6 against 68.5 and 69.8.

Latency, an internal test, and a bill that is six times larger

ModelMedian response (ms)p95 (ms)
Clef209.3238.6
Clef-flash38.8122.4
Jev524.1536.0

These are Cloudflare's measurements, not an independent benchmark, and latency on a vendor's own infrastructure is the easiest number for a vendor to control. The company adds a first-party use case rather than another chart: its Threat Intelligence team tested Clef on classifying website domains, and Clef took 2.2 seconds to fetch, render, and classify a domain against 4.7 seconds for its fastest general-purpose LLM, gpt-oss-120b, which returned only two classifications for the same input. That is an internal comparison with no outside check, and it says nothing about how often the classifications were correct.

Then the bill. The Register reports Clef at $0.24 per million tokens and describes that as nearly six times the price of Jev at $0.042 per million. Cloudflare's differentiators are speed and modality. Price is not among them.

Where the critics landed

The cost gap was the first thing the Hacker News thread did arithmetic on. At roughly 300 tokens per call, one commenter calculated, a million decisions cost about $12.60 on Jev and about $72 on Clef — a swing that no latency advantage fully answers. A second commenter waved off the scoreboard with "Public benchmarks are easy to cheat." A third argued that many models have claimed to beat Jev on the public benchmark and then stumbled on non-trivial tasks. None of those are measurements. They are the reasons the vendor's table cannot be the end of the argument, and at 614 points and 214 comments the thread treated them as the main event.

On r/LocalLLaMA (roughly 354 upvotes and about 107 comments), the sharper worry was about the comparison itself: commenters argued that Cloudflare's runs were measured against weaker open Jev variants rather than the strongest available build. That is community criticism, not an established fact, and Cloudflare has not answered it. It matters anyway, because the entire quality table rests on which version of Jev the vendor chose to put in the right-hand column.

What would settle it

Three things would move this past a launch post. An independent run of the quality table on a neutral harness, with the columns fixed by someone other than the vendor. A comparison against the current Jev build, since the community objection is precisely about version choice. And production telemetry — accuracy and latency in live traffic — which neither Cloudflare nor TypeSafe has published. Until then, Clef is an open-weight classifier with a credible speed story, a price six times the thing it is measured against, and a scoreboard its competitor did not get to write. Our earlier coverage of Jev's calibration fight is the other half of the same argument.

Related Articles

Scroll down

to load the next article