Cloudflare Clef: decision models at 9 cents per million tokens
Cloudflare's Clef and Clef-flash decision models are GA on Workers AI from $0.09 per million input tokens, Apache 2.0. Route and classify for pennies.
Cloudflare introduced Clef and Clef-flash on October 1: two open-source decision models hosted on Workers AI. They do not write paragraphs. They pick from options you define and return probabilities. Per Cloudflare's pricing page, Clef-flash costs $0.09 per million input tokens. Cloudflare joins a crowded week of vendors shipping models for the boring decisions inside an agent.
What actually happened
Per the Cloudflare announcement:
- Two sizes. Clef is built on a 27B Qwen base; Clef-flash on a 9B Qwen base. Both take text and images and have a 64k context window.
- Typed, bounded outputs. The models are non-autoregressive. They score a fixed set of choices and return probabilities instead of generating free text.
- Speed. Cloudflare reports 209.3 ms median latency for Clef and 38.8 ms for Clef-flash.
- Price. $0.24 per million input tokens for Clef and $0.09 for Clef-flash, with no output-token line, per the Workers AI pricing page. The free tier of 10,000 Neurons a day applies.
- Open weights. Both are on Hugging Face under Apache 2.0, so you can self-host if Workers AI is not your platform.
- Fine-tuning. Cloudflare also announced a reinforcement-learning product to tune decision models on your own data. Today it is run with Cloudflare's forward-deployed engineers; a self-serve version is "coming later." No price yet.
Cloudflare cites a 94.20% macro-F1 for Clef on BANKING77, a standard intent-classification benchmark. That is the vendor's number on a public dataset, not yours.
Why it matters for your business
Classification is most of your agent's calls. "Refund, sales lead or spam?" "Does this order need a human?" "Which tool runs next?" Each one is a pick from a list. Sending those to a frontier model means paying for reasoning you throw away. At $0.09 per million tokens, a million 100-token support tickets costs about $9 in input.
Probabilities give you a rule. A score lets you write: above 0.9, act; below, route to a person or a bigger model. We made the same case for AWS's Strands Decider and TypeSafe's Jev. Several vendors shipping the same pattern in one week tells you where agent architecture is going.
Apache 2.0 is the exit. Start on Workers AI, and if the pricing or the platform changes, the same weights run on your own GPU.
Key takeaways
- Clef (27B) and Clef-flash (9B) are GA decision models on Cloudflare Workers AI
- $0.24 and $0.09 per million input tokens; 209 ms and 39 ms median latency
- Apache 2.0 weights on Hugging Face, so you can self-host later
- RL fine-tuning exists but is engineer-assisted only, no public price
- Gate agent actions on the probability, and test on your own labeled data first
Paying frontier prices to sort tickets and route orders? We build agents that send the easy picks to a cheap decision model and only the hard cases to the big one. Estimate the savings or see how we build.
Sources: Cloudflare blog, Cloudflare Workers AI pricing.
- #cloudflare
- #clef
- #decision-models
- #workers-ai
- #ai-agents
Tommy Rush — Founder, Rush Commerce
Operator turned builder. 15+ years running operations — now shipping the systems businesses run on. More
Get The Rush Report weekly — one email, zero fluff.
Keep reading
OpenAI safety lead quits: 'culture is broken.' Plan a fallback
OpenAI safety lead David Robinson quit and called the company's culture broken. What a vendor's safety turmoil means for your AI stack, and the fallback plan to write now.
Read itCloudflare Web Search API: agent search on your AI Gateway bill
Cloudflare's Web Search API puts Exa, Linkup and Ceramic search behind AI Gateway at partner list price, no markup, with logs and one credit balance.
Read it