Cloudflare Just Shipped Two Decision Models. I Raced Them Against Jev (On OpenRouter)
A lot of AI calls these days don't really need to write anything. They just need to pick: which category, which passage, yes or no.
But most of us still use a chat model for that. We pay it to generate a paragraph, then parse out the one word we wanted.
Decision models skip the paragraph. You send the input and a list of allowed answers, and you get back a probability for every option. No text. One pass.
I've been using one of them, Jev from TypeSafe (~typesafe/jev-latest on OpenRouter, currently Jev 1.13), inside a small app that answers questions about long PDF manuals. Jev picks the best passages before the answer gets written.
Then on October 1 Cloudflare announced Clef in a post called Introducing Clef: our open-source decision models, and new RL fine-tuning platform (blog.cloudflare.com). Two models, actually:
- Clef: 27B, built on Qwen3.8-27B
- Clef-flash: 9B, built on Qwen3.5-9B
Both are Apache 2.0, with the weights on Hugging Face (huggingface.co/Cloudflare/clef), and both run on Workers AI (developers.cloudflare.com). Cloudflare calls Clef "fully Jev-API compatible", and both are on OpenRouter as cloudflare/clef and cloudflare/clef-flash (openrouter.ai).
So I tested both on their own, outside my app, and compared them with Jev on accuracy and speed.
What Cloudflare claims
The announcement and the model card make three big claims:
- faster than Jev. Clef at 209.3 ms median, Jev at 524.1 ms. Clef-flash at 38.8 ms
- as good or better on most tasks. For example 97.4% vs. 89.3% on CLINC150 intent classification. Jev wins on harder reasoning, like GPQA Diamond (78.3% vs. 48.0%)
- calibrated probabilities. Training used a Brier loss "to refine probability calibration"
All three are Cloudflare's own numbers, on Cloudflare's own hardware. Traictory pointed that out too: latency on your own infrastructure is "the easiest number for a vendor to control" (traictory.com).
I wanted to see what you actually get when you call it through OpenRouter.
What a decision request looks like
All three models sit behind the same OpenRouter endpoint, POST /api/alpha/decisions (OpenRouter's Jev docs). You send a state (what the model looks at) and questions (what it has to decide):
{
"model": "cloudflare/clef",
"state": {"message": "Tracking hasn't updated in five days."},
"questions": {
"intent": {
"type": "choice",
"instructions": "Which intent best describes the customer's request in state?",
"criteria": {
"billing": "Charges, invoices, refunds, payment methods, pricing",
"shipping": "Where a physical order is, delivery delays, returns of goods",
"bug": "Something in the product is broken or behaving wrongly"
}
}
}
}
Back comes the pick, a probability for every option, and a confidence number.
The nice part: switching models means changing one string.

Same answer from both. Note the confidence, though. We'll come back to it.
(The terminal exchanges in this post are reconstructed from the real session; the numbers are from the actual runs.)
The test
I kept it small and boring on purpose. Three tasks:
- intent: 40 short support messages I wrote and labelled by hand, each routed to one of 6 intents (billing, bug, account, feature, shipping, praise). Ten of them are tricky on purpose, like "Great app, but I was billed after my trial ended"
- passage: 30 questions about made-up companies, pick the passage that answers each one out of 10. Six of the 10 are traps: the company's other two passages, and three passages about its near-twin ("Corvid Labs" vs. "Corvid Systems")
- passage_all: 12 of those questions against all 120 passages at once, about 4.7K tokens per request
To keep the timing fair:
- every item goes to all three models, rotating who goes first
- the first call to each model is a warm-up and isn't counted
- I report the median and the slow end (p90), not the average
The results
I ran it twice. First Jev against Clef, then, a bit later, all three together. This is the second run, so all three models faced the same network at the same time:

The Errors column is Clef: in each task, one call failed with HTTP 429 and "Capacity temporarily exceeded" from Workers AI, and it counts as wrong. Every Clef answer that actually came back was correct, except one.
That one was the message every model got wrong: "Is there a way to log in with Google?" I labelled it feature, all three said account. Honestly, I'd label it differently myself today.
So on answers, Clef ties with Jev. The first run said the same: 39/40, 30/30 and 12/12 for Clef, against 38/40, 30/30 and 12/12 for Jev.
Clef-flash is a different story.
Of its 5 wrong passages, 4 were the same mistake: right kind of fact, wrong company. Asked what Rowan Systems sells, it answered with what Rowan Labs sells. Exactly the near-twin trap the test was built to catch. The bigger models didn't fall for it once.
Now speed.

Jev was the fastest on every task, in both runs.
Clef was 1.7–2.4× slower on small requests and 3.4× slower on the long one, in both runs. Clef-flash landed in between: 1.2–1.3× slower than Jev on small requests, 2.2× on the long one.
That's the opposite of Cloudflare's chart. They measured Clef at 209 ms, Clef-flash at 38.8 ms and Jev at 524 ms. I never saw either Clef model beat Jev.
My numbers are the full round trip from my machine through OpenRouter, which is what your app sees. Theirs are on their own setup. Both can be true. Only one of them is what you get today.
(The second run was slower across the board, Jev included, so compare models within a run, not across runs.)
The price surprise
Clef's listing says $0.24 per million input tokens and $0 for output, the same as on Workers AI (openrouter.ai, developers.cloudflare.com). Since a decision model doesn't write any output, that sounds cheap.
Clef-flash is listed at $0.09 (openrouter.ai).
Neither is cheap next to Jev. Jev is listed at $0.042 per million input tokens, also with free output (openrouter.ai). The usage field in my responses matched all three prices. I'm not the only one who noticed: The Register and Hacker News called out the same 6× gap on launch week (traictory.com).
The Clef models count fewer tokens for the same text, but pay more per token: about 6× for Clef and 2× for Clef-flash.
Per call, that made Clef 3.4–6× more expensive than Jev and Clef-flash 1.3–2.3× more.
To be fair: both runs together, 410 counted calls plus warm-ups, cost about 5 cents. You won't feel this at small scale. You will at a million calls a day.
Does it know when it's wrong?
This is the part that actually matters for my app.
Jev doesn't just pick. It says how sure it is, and I use that. If the top passage comes back with low confidence, that's a hint the manual may not answer the question at all.

Jev: 0.99 when right, 0.71 when wrong. A clear gap.
Clef: 0.81 when right, 0.74 when wrong. Almost no gap. And on the long task it got every answer it returned right while saying it was about 41% sure.
Clef-flash is interesting. On intents its gap was the biggest of all three: 0.81 vs. 0.43. But on the long task it flipped, and it was less sure when right than when wrong.
A confidence number that's the same whether you're right or wrong isn't a confidence number. It's noise.
That's odd, given Cloudflare trained it specifically for calibration. Maybe it shows on bigger and harder sets. On mine, it didn't.
So even with equal accuracy, Clef isn't a drop-in swap for anything that acts on the confidence. And Clef-flash's confidence only works some of the time.
The Limitations
- small, hand-made sets. 40, 30 and 12 items. Good enough to spot a 2× speed gap, not to rank models that are this close on accuracy
- my labels, my wording. The tasks are easy on purpose, so Jev and Clef both hit the ceiling
- two runs, one location, one day. Clef is four days old. Latency on a brand-new endpoint often improves, and the 429s suggest Cloudflare is still adding capacity
- no retries. A 429 counts as a miss. In production you'd retry, which fixes the answer but makes Clef even slower
- the 2K-token note isn't tested properly. Clef's OpenRouter page says Workers AI reads only about the first 2K tokens of long text in
state(openrouter.ai). Cloudflare's own docs just say long text is "truncated to fit the model's token limit" (developers.cloudflare.com). My long request had 4.7K tokens of options and a shortstate, and Clef read it fine. I did not try a longstate, which is what that note is actually about - Clef has things I didn't test. It reads images and video, its context is 64K against Jev's 32K, and the weights are open, so you can run it yourself. If you want to self-host or you're already on Cloudflare, that alone may decide it
The Bottom Line
Clef is a real decision model. It picks as well as Jev on simple tasks, and open weights are a genuine plus.
But on OpenRouter today it's slower, it gets slower as the input grows, it costs more per call, it sometimes runs out of capacity, and its confidence doesn't tell you much.
Clef-flash is closer to Jev on speed and price. But it's still slower, still pricier, and it mixes up look-alike names.
For my app, Jev stays. I'll run the same test again in a month.