Skip to content
Rush Commerce
AI & Automation3 min read

Cloudflare Clef: decision models at 9 cents per million tokens

Cloudflare's Clef and Clef-flash decision models are GA on Workers AI from $0.09 per million input tokens, Apache 2.0. Route and classify for pennies.

Cloudflare introduced Clef and Clef-flash on October 1: two open-source decision models hosted on Workers AI. They do not write paragraphs. They pick from options you define and return probabilities. Per Cloudflare's pricing page, Clef-flash costs $0.09 per million input tokens. Cloudflare joins a crowded week of vendors shipping models for the boring decisions inside an agent.

What actually happened

Per the Cloudflare announcement:

  • Two sizes. Clef is built on a 27B Qwen base; Clef-flash on a 9B Qwen base. Both take text and images and have a 64k context window.
  • Typed, bounded outputs. The models are non-autoregressive. They score a fixed set of choices and return probabilities instead of generating free text.
  • Speed. Cloudflare reports 209.3 ms median latency for Clef and 38.8 ms for Clef-flash.
  • Price. $0.24 per million input tokens for Clef and $0.09 for Clef-flash, with no output-token line, per the Workers AI pricing page. The free tier of 10,000 Neurons a day applies.
  • Open weights. Both are on Hugging Face under Apache 2.0, so you can self-host if Workers AI is not your platform.
  • Fine-tuning. Cloudflare also announced a reinforcement-learning product to tune decision models on your own data. Today it is run with Cloudflare's forward-deployed engineers; a self-serve version is "coming later." No price yet.

Cloudflare cites a 94.20% macro-F1 for Clef on BANKING77, a standard intent-classification benchmark. That is the vendor's number on a public dataset, not yours.

Why it matters for your business

Classification is most of your agent's calls. "Refund, sales lead or spam?" "Does this order need a human?" "Which tool runs next?" Each one is a pick from a list. Sending those to a frontier model means paying for reasoning you throw away. At $0.09 per million tokens, a million 100-token support tickets costs about $9 in input.

Probabilities give you a rule. A score lets you write: above 0.9, act; below, route to a person or a bigger model. We made the same case for AWS's Strands Decider and TypeSafe's Jev. Several vendors shipping the same pattern in one week tells you where agent architecture is going.

Apache 2.0 is the exit. Start on Workers AI, and if the pricing or the platform changes, the same weights run on your own GPU.

Key takeaways

  • Clef (27B) and Clef-flash (9B) are GA decision models on Cloudflare Workers AI
  • $0.24 and $0.09 per million input tokens; 209 ms and 39 ms median latency
  • Apache 2.0 weights on Hugging Face, so you can self-host later
  • RL fine-tuning exists but is engineer-assisted only, no public price
  • Gate agent actions on the probability, and test on your own labeled data first

Paying frontier prices to sort tickets and route orders? We build agents that send the easy picks to a cheap decision model and only the hard cases to the big one. Estimate the savings or see how we build.

Sources: Cloudflare blog, Cloudflare Workers AI pricing.

  • #cloudflare
  • #clef
  • #decision-models
  • #workers-ai
  • #ai-agents
TR

Tommy Rush — Founder, Rush Commerce

Operator turned builder. 15+ years running operations — now shipping the systems businesses run on. More

Get The Rush Report weekly — one email, zero fluff.