Cloudflare Clef open decision models graphic with price and license

Cloudflare Clef Explained: Open Decision Models That Score Answers, Not Write Essays

Cloudflare Clef is a pair of open-weight “decision models” that answer yes/no, multiple-choice and ranking questions with probabilities instead of paragraphs. Cloudflare announced them on October 1, 2026, and you can use them on Workers AI or download them from Hugging Face. If you build AI agents that need fast, cheap, structured answers, they are worth a look.

Here are the facts that matter most.

  • Two models: Clef (27B) and Clef-flash (9B), announced October 1, 2026.
  • Price: $0.24 per million input tokens for Clef, $0.09 for Clef-flash, with no output charge.
  • License: Apache 2.0 weights, but the training data is not public.
  • Compatible with the API of TypeSafe AI’s Jev, so it can drop into existing setups.

That price gap and the self-reported benchmarks are the two things to read carefully before you switch anything.

What Cloudflare Clef Actually Does

Normal chatbots write text one token at a time. A decision model skips that. You send it some context and a list of questions, and it scores the answer options in a single pass.

According to The Register, both models handle bounded questions such as yes/no, multiple choice and rankings. Clef takes text and images, and The Register says video too, with a 64k-token context window.

Think of tasks like “Is this ticket about billing?” or “Which of these five tools should the agent call next?” You get probabilities back, not a paragraph you have to parse. That is useful when an agent makes dozens of small choices per task. We covered a similar idea from OpenAI in our OpenAI Decisions API guide.

Because nothing is generated, there is no long answer to wait for or pay for. That is where the speed and the lower price come from. It also means you cannot ask Clef to explain itself, so keep a regular chat model around for tasks that need written reasoning.

Clef vs Clef-flash: Which One Fits Your Job

Clef is the bigger, more accurate model. Clef-flash trades some accuracy for much lower latency and cost. Here is how they compare.

Detail Clef Clef-flash
Size 27B parameters 9B parameters
Built on Qwen3.8-27B Qwen3.5-9B
Price per 1M input tokens $0.24 $0.09
Median decision time 209.3 ms 38.8 ms
GPU memory to self-host About 85 GB About 41 GB

The prices and latency come from Developers Digest, and the memory figures from The Register, which assumes one request at a time and a full 64k context. Neither model charges for output, because nothing is generated.

If your decisions are simple and speed matters, start with Clef-flash. Move to Clef only if accuracy slips on your own data.

How Clef Compares With Jev, and Why to Be Careful

Jev, from TypeSafe AI, drew attention about two weeks before Clef arrived. Cloudflare built its API to match Jev, and published its own comparison.

Test Clef Jev
BFCL (tool calling, exact match) 98.5 95.8
BANKING77 (intent, macro-F1) 94.2 79.7
GPQA Diamond (hard science questions) 48.0 78.3
MMLU-Pro (broad knowledge) 65.9 82.7

So Clef wins on classification and tool-choice tasks, while Jev wins on hard reasoning and knowledge tests. These are Cloudflare’s own runs, as reported by Developers Digest, and they have not yet been reproduced on the official Decision Index.

Cost is the other catch. The Register puts Jev at $0.042 per million tokens, so Clef costs nearly six times as much. You are paying for accuracy on certain tasks, not a blanket upgrade.

How to Try Cloudflare Clef on Workers AI

You call it over a normal REST endpoint with a bearer token. Per Developers Digest, the path looks like this:

/client/v4/accounts/{account_id}/ai/run/@cf/cloudflare/clef-flash

Swap in clef for the larger model. You send a state (text or JSON) plus a list of questions, each typed as true/false, choice or score. The response gives a probability for each option.

Keep these request limits in mind before you design around it.

  • Up to 64 questions per request.
  • Up to 4 images per request, 4 MiB each.
  • A 13 MiB request body cap.
  • States longer than 64k tokens get truncated.

Those caps decide how you batch work, so test with your real inputs first. If you run agents on your own hardware, our look at Nvidia’s open agent safety platform covers the guardrail side of the same problem.

One practical tip: log every decision with its probabilities for the first few weeks. A low-confidence answer is a good signal to send that case to a person or a larger model, and the logs show you where Clef-flash is good enough.

Frequently Asked Questions

What is a decision model?

It is an AI model that scores a fixed set of answer options instead of writing free text. You get probabilities for each choice in one pass.

Is Cloudflare Clef open source?

Not fully. The weights are Apache 2.0 and downloadable from Hugging Face. But Cloudflare’s Michelle Chen confirmed to The Register that the training datasets are not public, so “open weights” is the safer term.

How much does Cloudflare Clef cost?

Clef is $0.24 per million input tokens and Clef-flash is $0.09. There is no output charge, since the models generate no text.

Can I run Clef on my own GPU?

Yes, if you have enough memory. The Register says Clef-flash needs at least 41 GB of VRAM and Clef needs 85 GB, with a 64k context and one request at a time.

Who Should Switch, and Who Should Wait

If you already use a Jev-compatible setup and your tasks are classification or tool routing, Clef-flash is an easy, cheap test because the API matches. If you need strong reasoning scores or want audited benchmarks, wait for independent results on the Decision Index. Either way, run it on your own data first, because vendor numbers rarely match real workloads.


Comments

Leave a Reply

Your email address will not be published. Required fields are marked *