Cloudflare shipped two decision models on October 1, and the weights are on Hugging Face under Apache 2.0. Clef is a 27 billion parameter model. Clef-flash is 9 billion. Both are hosted on Workers AI, both answer the same API shape as TypeSafe's Jev, and Cloudflare's own numbers put Clef ahead of Jev on 7 of 10 benchmarks. Jev entered early access on September 15. Sixteen days later the category it created has a clone with open weights from one of the largest networks on the internet.

How a decision model answers

A decision model takes a block of state (a support ticket, a log line, an HTTP request, a screenshot) and a list of typed questions, and returns a probability for each allowed answer. No prose. No tokens generated one at a time. TypeSafe coined the term "System One model" for this in September, borrowing from Kahneman, and priced Jev at $0.042 per million input tokens with output unmetered.

Clef supports three question types. A yes/no question (Cloudflare calls it "noul") returns one probability. A choice question with 2 to 255 options returns a probability per option. A score question with 2 to 10 levels returns a weighted score that can land between levels. You can pack up to 64 questions against one state into a single request, and the context window is 64k tokens.

The speed comes from how the answer gets computed. The Qwen backbone reads the state and the questions in one prefill pass. A separate scoring head then looks at every allowed answer in parallel and runs a softmax over them. There is no decode loop. Cloudflare's announcement reports Clef at 209 ms median and 239 ms at p95 across 43 benchmark runs, and Clef-flash at 39 ms median and 122 ms at p95. Cloudflare's measurement of Jev in the same test setup was 524 ms median. TypeSafe's own docs claim 70 to 500 ms, so treat the comparison across vendors as a rough shape and not a fact.

Built in 16 days on Qwen

Cloudflare didn't pretrain anything. Clef is Qwen 3.8-27B with the base weights frozen, plus rank 256 LoRA adapters trained with label smoothing, cross-entropy loss, and a Brier loss term for calibration. Clef-flash is the same recipe on Qwen 3.5-9B. On top sits a reinforcement learning stage Cloudflare calls RLCD, which gives partial credit when a score answer lands one level off from the label. The training data is synthetic and not published.

That last sentence is the one the Hacker News thread fixed on. One commenter put it as "weights are not source." The license on the weights is Apache 2.0 and the Qwen base is open, but you cannot reproduce Clef from what Cloudflare released. Cloudflare's blog says open source. The Register made the same objection I'm making. Open weights is the honest label and I'll use it.

The turnaround also answers a question several people asked in the thread: is this hard? It is not. An instruction tuned base model plus a classification head plus a decent synthetic dataset gets you a decision model. OpenAI announced a Decisions API built on a trimmed GPT-6 Luna at DevDay on September 29, and the thread names smaller clones. Jev's architecture is still secret. TypeSafe's moat, if it has one, is the training data and the calibration work. The idea itself took Cloudflare 16 days to copy.

Where Clef wins and where Jev wins

Cloudflare published scores on TypeSafe's own Decision Index. Clef leads on BANKING77 at 94.2 macro-F1 against Jev's 79.7, on CLINC150 with out of scope detection at 97.4 against 89.3, and on BFCL function calling at 98.5 against 95.8. Jev leads where the question needs reasoning: GPQA Diamond 78.3 to Clef's 48.0, MMLU-Pro 82.7 to 65.9, When2Call 81.0 to 72.4.

Read that split as a product spec. Clef is the better intent classifier and the better tool router. Jev is the better judge when the right answer takes a chain of thought to reach. If your agent's gating question is "which of these 40 tools runs next," use Clef. If it is "does this proof hold," use neither. Call a reasoning model and pay for it.

Two caveats. Every number above is Cloudflare's run, and nobody outside Cloudflare has reproduced them yet. And Flavio Copes, who spent launch day hammering the API, saw image requests take 13 to 30 seconds and had to shrink a 1.3 MB screenshot to 140 KB before the call succeeded. The vision encoder exists. It is not ready for a synchronous path.

Pricing, VRAM, and where I'd use it

Hosted on Workers AI, Clef is $0.24 per million input tokens and Clef-flash is $0.09. Output is not billed. Jev is $0.042. So the big model is 5.7 times Jev's price and the small one is 2.1 times. One million decisions against a 2,000 token state is 2 billion input tokens, which is $480 on Clef, $180 on Clef-flash, and $84 on Jev.

Those are all small numbers next to a general LLM. GPT-5.6 Terra bills $2.00 per million input tokens and $12.00 on output. The pricing fight between Clef and Jev is a rounding error for most teams. What Cloudflare's premium buys you is the right to download the weights and leave.

Leaving has a hardware cost. The Register reports Clef-flash needs 41 GB of VRAM and Clef needs 85 GB, both at single concurrency with the full 64k context. Clef-flash fits on one 48 GB card, or on an 80 GB H100 with room to batch. Clef does not fit on a single H100 at full context. That puts the full model out of reach for a homelab and into one node in the rack for everyone else. At the volumes where self hosting pays back, fine. For a startup making 50,000 decisions a day, pay the $0.24.

Cloudflare is running Clef on itself. The threat intelligence team classifies domains with it in 2.2 seconds per batch where gpt-oss-120b took 4.7. Trust and safety uses it to triage abuse reports. Support uses it to route tickets. Bot management uses it to separate good crawlers from bad. Those are the same four jobs every ops team has, and today each of them is a regex pile, a rules engine nobody wants to touch, or a full LLM call that costs an order of magnitude more and takes two seconds.

My rule for the next six months: any place in a pipeline where a model picks from a closed list is a decision model job. Routing. Severity scoring. Escalation gates. Alert dedup. Use hosted Clef-flash to find out whether the model can answer your question at all, and keep the probabilities. A 0.5 answer usually means your question is ambiguous. Fix the question before you blame the model. Copes made that point and it matches what I've seen from classifiers since long before anyone called them decision models.

The RL fine-tuning platform that shipped alongside is a service today and a product later. Right now it means Cloudflare's forward deployed engineers capture your traffic through AI Gateway, generate rollouts on Workers AI, run the sandbox in Containers, and push the new adapter back with bring your own model. Self serve has no date. If you have an enterprise account, call them. If you don't, the Apache 2.0 weights plus a LoRA script get you most of the way on your own GPU.

Jev had the category to itself for 16 days. Cloudflare and OpenAI now sell the same shape of API. My prediction: by the end of 2026 the hosted price for a 9B class decision model is under five cents per million tokens from every vendor, and the two differences left are calibration quality and whether you can run the thing yourself. TypeSafe has the first. Cloudflare has the second. Nobody has both.