Cloudflare Clef Decision Models: Open‑Source, Vision‑Enabled, and RL‑Fine‑Tunable
TL;DR
Clef and Clef‑flash are Cloudflare‑trained, open‑source decision models that beat Typesafe’s Jev on the Jev Decision Index, run up to 10× faster, support image inputs, and can be fine‑tuned via a new RL‑based service.
What is a decision model?
A decision model takes structured inputs (e.g., a support ticket, a URL) and returns typed classifications with associated probabilities. The output can be directly consumed by code to route tickets, trigger escalations, or defer to a human. Unlike general‑purpose LLMs, decision models are designed for deterministic‑looking, bounded outputs and low latency.
"A decision model makes classifications to help agents decide how to act, based on certain probabilities." – Cloudflare blog
Example use case at Cloudflare
- Input: a website domain fetched with Browser Run.
- Output: probabilities such as 95 % fashion, 85 % e‑commerce, <1 % phishing.
- Latency: 2.2 s vs 4.7 s for the fastest LLM (gpt‑oss‑120b) on the same workflow.
How Clef differs from other decision models
| Feature | Clef | Clef‑flash | Jev (Typesafe) |
|---|---|---|---|
| Base model | Qwen‑3.8‑27B (frozen) | Qwen‑3.8‑9B (frozen) | Proprietary |
| Vision encoder | ✅ (image classification) | ✅ | ❌ |
| Context window | 64 k tokens | 64 k tokens | 32 k tokens |
| Latency (median) | 209 ms | 38.8 ms | 524 ms |
| p95 latency | 238 ms | 122 ms | 536 ms |
| Benchmark leader | Highest scores on the Jev Decision Index across 43 evals | Near‑state‑of‑the‑art on speed‑critical tasks | Competitive but generally lower |
Clef’s vision encoder makes it the first non‑generative decision model that can classify visual content, a capability Jev lacks. The larger context window lets users provide more state (e.g., longer logs) without truncation.
Benchmark performance
Clef leads the public Jev Decision Index leaderboard. Selected scores (higher is better):
- BFCL case exact – 98.47 % (Clef) vs 95.75 % (Jev)
- API‑Bank accuracy – 93.11 % vs 88.19 %
- BANKING77 macro‑F1 – 94.20 % vs 79.74 %
On the Typesafe workflow suite, Clef‑flash outperforms Jev on invoice processing (64.7 % vs 61.8 %) and agent‑trace observability (71.6 % vs 68.5 %).
"These models are smarter, faster, and fully Jev‑API compatible, so you can experiment with these hosted models easily." – Cloudflare blog
Architecture that enables speed
- Two‑stage attention routing – The model first runs a prefill‑only pass of Qwen, then scores each schema choice in parallel. No autoregressive token generation is required.
- Option‑specific evidence extraction – Each possible answer extracts relevant context, cross‑attends with other fields, and finally receives a probability score.
- Low‑rank adapters – Rank‑256 adapters are trained on top of frozen Qwen weights, allowing rapid post‑training without full model fine‑tuning.
- Loss functions – Label‑smoothed cross‑entropy for correct schema outputs combined with a Brier loss for calibrated probabilities.
- Reinforcement Learning for Calibrated Decisions (RLCD) – Provides partial credit for near‑misses, penalises distribution shift, and improves both accuracy and calibration.
The non‑autoregressive design eliminates the token‑by‑token generation bottleneck, delivering up to 10× lower latency than comparable LLM‑based classifiers.
Hosting on Workers AI
Clef models are deployed on Cloudflare’s edge‑located GPUs via Workers AI, giving:
- Sub‑millisecond network round‑trip times.
- Ability to place the model in the hot path of request handling.
- Simple HTTP API compatible with the Jev schema.
Example curl request (excerpt):
curl https://api.cloudflare.com/client/v4/accounts/$CLOUDFLARE_ACCOUNT_ID/ai/run/@cf/cloudflare/clef \
-X POST -H "Authorization: Bearer $CLOUDFLARE_AUTH_TOKEN" \
-d '{
"model": "clef",
"state": "Checkout has been failing for every customer for the last hour.",
"questions": { ... }
}'
Open‑source release
- Repository: https://huggingface.co/Cloudflare/clef (Apache 2.0)
- Weights are provided for both Clef and Clef‑flash.
- The data and training pipeline are not open‑sourced; only the final model weights are released under a permissive license.
"Open‑source weights, not open source. The weights have permissive licensing, but the data and training pipeline are not published to reproduce them from their proprietary Qwen starting points." – HN comment by buildbuildbuild
Fine‑tuning via RL service
Cloudflare introduces a reinforcement‑learning fine‑tuning platform that lets customers adapt Clef to domain‑specific tasks (e.g., Trust & Safety triage, bot classification). The workflow leverages existing Cloudflare primitives:
- AI Gateway – Captures request/response pairs to build a dataset.
- Workers AI – Generates rollouts against the base Clef model.
- Containers – Provides an RL sandbox for scoring and replay.
- Trainer – Updates model weights using the RLCD objective.
- BYO Model deployment – Redeploys the fine‑tuned model on Workers AI.
The service is initially a hands‑on partner offering with the Forward‑Deployed Engineer (FDE) team, with a self‑serve platform planned for later.
Community reaction on Hacker News
| Comment theme | Representative quote |
|---|---|
| Performance vs cost | "Pricing is $0.24 / million input tokens, ~6× higher than Jev. Clef‑flash is $0.09 / M, more competitive." – ssiddharth |
| Open‑source concerns | "Open weights, not open source. Data and training pipeline are not published." – buildbuildbuild |
| Vision capability | "It allows image input. Nice!" – swingboy |
| Speed vs accuracy trade‑off | "Clef‑flash is faster but sometimes over‑escalates; Clef is more accurate but slower." – agrippanux |
| Real‑world utility | "If Clef can decide domain categorisation in 2 s, I shouldn’t have to file a support ticket for uncategorised sites." – johnbatch |
| Skepticism about novelty | "How did many teams build decision models within days of Jev’s release? Is the concept already mature?" – warkdarrior |
| Calibration doubts | "No mention of calibration. Is it just another LLM fine‑tune?" – zwaps |
Overall sentiment is positive about the open‑source release and speed gains, but reviewers note higher pricing, lack of reproducible training pipelines, and the need for fine‑tuning to reach production‑grade accuracy.
When to use Clef vs a full LLM
- Use Clef when you need:
- Structured, probability‑based decisions.
- Low latency (sub‑200 ms) at edge locations.
- Vision‑enabled classification.
- Deterministic‑looking outputs for downstream automation.
- Use a generative LLM when you need:
- Free‑form text generation, reasoning, or tool‑call orchestration.
- Complex multi‑turn dialogues.
- Situations where the overhead of a few hundred milliseconds is acceptable.
Getting started
- Try the hosted endpoint – Follow the API example above.
- Download the weights –
git clone https://huggingface.co/Cloudflare/clef. - Run locally – Use the standard
transformersorvLLMpipelines (the model expects the Jev‑compatible schema). - Explore the live benchmark demo – https://clef-evals.workers-ai-mle.workers.dev/.
- Contact Cloudflare for fine‑tuning – Fill out the interest form linked in the blog post.
Future outlook
Clef demonstrates that decision‑focused models can be both high‑quality and edge‑ready, carving a niche between heavyweight LLMs and traditional classifiers. The upcoming self‑serve RL fine‑tuning platform could further accelerate adoption by allowing organisations to tailor the model to proprietary domains without leaving Cloudflare’s infrastructure.
If you’re building agentic workflows that require rapid, structured decisions—especially with visual inputs—Clef and Clef‑flash provide a compelling, open‑source alternative to Jev and generic LLMs.
Sources
Related
- Dispatch
- Dispatch
- Dispatch
- Dispatch
- Dispatch