Introducing Clef: our open-source choice fashions, and new RL fine-tuning platform

Over the previous couple of weeks, there was a lot of buzz round choice fashions comparable to Typesafe AI’s Jev System One mannequin. While classifier fashions have been round for a while, Jev introduces a brand new choice mannequin idea into the globe of AI — a mannequin that produces bounded structured outputs cheaply, shortly and constantly that may be added right into a workflow when a choice is required. These fashions are succesful sufficient to work over any set of inputs with out continually retraining the mannequin to include new classification classes. This contrasts with the globe of Large Language Models (LLMs), that are largely non-deterministic, however are open-ended sufficient to motive and generate textual content and gear requires agentic workloads.
Today, we’re releasing two Cloudflare-trained choice fashions, Clef and Clef-flash, hosted on Workers AI. Clef is presently the chief when evaluated towards the Jev Decision Index, you possibly can view full outcomes on the live benchmark demo site. These fashions are smarter, sooner, and absolutely Jev-API suitable, so you possibly can experiment with these hosted fashions simply. We’re absolutely open-sourcing these models on Hugging Face beneath an Apache 2.0 license so that you can run regionally and experiment with yourselves.

Lastly, we’re excited to debut our new reinforcement studying (RL) product, which permits clients to fine-tune Clef to go well with their use instances as nicely.
What is a choice mannequin?
A call mannequin makes classifications to assist brokers resolve the way to act, based mostly on sure chances. For instance, you possibly can cross in a buyer assist message (inputs) and ask whether it is pressing and which workforce ought to deal with it. A call mannequin will return typed solutions with chances (outputs), which your code can use to route the ticket, set off an escalation, or defer to a human. This signifies that a human doesn’t essentially should be within the loop for agentic selections anymore — brokers can programmatically collect context, make selections, and take actions on duties, or defer to a human when wanted.

Specifically at Cloudflare, we’ve been testing our new Clef mannequin on our Threat Intelligence workforce to assist us classify web site domains. By giving a site to Clef (with Browser Run) it may well shortly establish classes that the area falls beneath — for instance, it would classify a site with a 95% probability it’s a style web site, 85% ecommerce, <1% phishing, and many others. This classification took our Clef mannequin 2.2s to fetch, render, and classify the web site. In distinction, our quickest basic LLM gpt-oss-120b took 4.7s in the identical workflow, and solely returned two classifications. As a consumer, you possibly can think about how a 2x financial savings in latency and outcomes will help us enhance our risk intelligence workflows and be sooner in figuring out malicious or reliable domains. Generalize this to any use case the place it’s essential make fast programmatic selections, and also you unlock highly effective agentic workflows which can be capable of autonomously resolve, motive, and execute.
In music concept, a clef is an emblem positioned at first of a musical workers that assigns particular pitch names to the strains and areas. A call mannequin is analogous to a music clef as a result of it helps outline the area of the context and the next notes (actions) that observe it. We selected Clef because the title of our household of choice fashions, because it serves related functions, and the CF hearkens to Cloudflare.
How is Clef totally different from different choice fashions?
Although the market is getting more and more saturated with choice fashions, Clef has some distinctive properties that make us excited to launch it to the general public. First, it has a imaginative and prescient encoder so it’s in a position to absorb photographs and classify visible content material. This is totally different from Jev, which solely does textual content classification at the moment. Secondly, our mannequin has a 64k context window (in comparison with Jev’s 32k), which permits customers to squeeze extra enter state for the mannequin to categorise towards.
Third, our mannequin is correct and highly effective, scoring competitively towards different choice fashions available on the market throughout numerous high quality benchmarks. We shortlisted some evaluations beneath which can be vital for decision-making as outlined by the Jev Decision Index and scored a number of the extra fashionable fashions available on the market for it. Check out the desk beneath for benchmarks, or view the scores on our live decision index demo site:
|
Benchmark |
DiffusionGemma Jev |
|||||
|
BFCL · case actual |
98.47 |
98.76 |
95.75 |
96.52 |
94.51 |
38.13 |
|
ToolRet · nDCG@10 |
69.19 |
66.43 |
65.28 |
61.21 |
64.26 |
12.69 |
|
API-Bank · accuracy |
91.93 |
93.11 |
88.19 |
83.66 |
56.30 |
11.41 |
|
Home home equipment · case actual |
82.95 |
97.73 |
52.27 |
42.05 |
25.00 |
0.00 |
|
When2Call · accuracy |
72.37 |
65.58 |
80.97 |
75.44 |
49.62 |
11.94 |
|
BANKING77 · macro-F1 |
94.20 |
90.93 |
79.74 |
74.28 |
84.83 |
14.29 |
|
CLINC150+OOS · macro-F1 |
97.43 |
66.77 |
89.27 |
83.49 |
79.03 |
3.19 |
|
BRIGHT · nDCG@10 |
45.91 |
39.26 |
47.52 |
42.94 |
38.53 |
19.90 |
|
Amazon ESCI · macro-F1 |
57.48 |
57.39 |
55.21 |
53.37 |
49.22 |
24.40 |
|
PhishNChips · accuracy |
79.60 |
75.05 |
62.55 |
85.35 |
50.75 |
50.15 |
We additionally ran benchmarks throughout Typesafe’s own eval suite and our Clef fashions fared nicely, beating Jev in 3 out of 4 areas. Notably, our Clef-flash performs exceptionally nicely, given how a lot sooner it’s.
|
Workflow |
|||
|
Invoice processing |
64.7 |
57.1 |
61.8 |
|
Customer service |
76.3 |
77 |
76.0 |
|
Security incidents |
62.9 |
61.7 |
61.7 |
|
Agent hint observability |
68.5 |
69.8 |
71.6 |
Across the 43 eval benchmarks that we ran, our Clef fashions beat the choice fashions on latency (aside from Laya which could be very quick however trades off high quality within the benchmarks above):
|
Benchmark |
DiffusionGemma Jev |
|||||
|
Median latency · ms |
209.3 |
38.8 |
524.1 |
84.4 |
51.4 |
5.8 |
|
p95 latency · ms |
238.6 |
122.4 |
536.0 |
211.2 |
187.9 |
222.5 |
On high of the latency advantages from the mannequin itself, our Clef fashions are hosted on Workers AI. Because they’re hosted on Cloudflare’s roads, we’re capable of benefit from our GPUs on the edge, resulting in low community latency and sooner selections. This signifies that you could possibly put Clef into the recent path for brokers to make selections and mix that with considered one of our LLMs on Workers AI to take motion.
curl https://api.cloudflare.com/client/v4/accounts/$CLOUDFLARE_ACCOUNT_ID/ai/run/@cf/cloudflare/clef
-X POST
-H "Authorization: Bearer $CLOUDFLARE_AUTH_TOKEN"
-d '{
"mannequin": "clef",
"state": "Checkout has been failing for each buyer for the final hour.",
"questions": {
"pressing": { "kind": "noul", "directions": "Is this assist request pressing?" },
"workforce": {
"kind": "selection",
"directions": "Which workforce ought to deal with this request?",
"standards": {
"billing": "Payments, invoices, and refunds",
"technical": "Outages, errors, and configuration",
"gross sales": "Plans and upgrades"
}
},
"severity": {
"kind": "rating",
"directions": "How extreme is the shopper influence?",
"standards": ["No impact", "Minor", "Major", "Critical"]
}
}
}'
Clef additionally produces strictly typed outputs just like Jev and is absolutely API-compatible, so you can also make the swap extraordinarily simply. The bigger Clef mannequin is your extra highly effective precision mannequin, whereas the Clef-Flash mannequin is nice for latency-critical selections. The fashions are enterprise-ready with our assure that we don’t learn, retailer, or practice in your requests or responses (until you wish to use our fine-tuning product, which we go into beneath). You can get began with the Clef fashions at the moment, beginning with our developer documentation or mess around with the open-source model on the Hugging Face repo.
If you’d like assist tuning Clef for a particular workload, we’re additionally providing fine-tuning companies — first as a hands-on companion with our forward-deployed engineer (FDE) workforce, after which later as a self-serve fine-tuning platform for purchasers to coach and redeploy the mannequin onto Cloudflare.
How we educated Clef
In the identical week that Jev got here out, we posted about some experiments we had with our personal homegrown choice mannequin. Our demo goes into how we tailored the DiffusionGemma mannequin to output deterministic chances by exposing the logprobs which can be generated by a big language mannequin. Our preliminary method constructed upon impartial analysis by Matt Mastracci, who has been energetic within the machine studying (ML) neighborhood with sharing new concepts and pull requests to vLLM inference engine to make DiffusionGemma assist stronger.
Clef builds upon this idea, however makes use of a unique base mannequin because the spine. We presently use Qwen as the bottom mannequin and post-trained it to go well with choice mannequin use instances. During inference, Clef makes use of Qwen for a prefill-only cross, then scores the legitimate schema decisions in parallel. The choice step is non-autoregressive, so there’s no intermediate textual content to generate token by token, making Clef considerably sooner than autoregressive LLMs. Rather than producing intermediate textual content to provide structured solutions, Clef and Clef-flash derive schema decisions instantly from inside spine representations. This method depends on a specialised two-stage consideration routing course of: each legitimate selection extracts context related to the immediate, permitting particular person subject parameters to cross-attend with different fields and again to the unique payload previous to scoring. By leveraging a lexical prior, the mannequin preserves semantic intent throughout choices. Ultimately, the structure unites option-specific proof routing, joint cross-field consideration, and schema-bound scoring.
By freezing Qwen3.8-27B for Clef and Qwen3.5-9B for Clef-flash, we collectively optimized the routing head alongside rank-256 low-rank adapters. Our post-training makes use of label-smoothed cross-entropy for legitimate schema outputs paired with a Brier loss to refine likelihood calibration. This coaching leverages our personal inside artificial datasets permutating subject orders, prompts, and schema constructions. We additionally developed Reinforcement Learning for Calibrated Decisions (RLCD) to function a secondary optimization goal, granting partial credit score to adjoining ordinal decisions, rewarding absolutely exact document outputs, and making use of a reference penalty to forestall distribution shift, giving us higher accuracy and generalization.
This signifies that we have been capable of obtain a number of novel issues with Clef: we improved accuracy of the mannequin in classification, constrained it to output solely chances as an alternative of textual content technology, and made it sooner than Jev and the bottom Qwen fashions.
How fine-tuning can prolong the capabilities of Clef
We heard a whole lot of inside use instances that required fine-tuning our Clef mannequin to be constructed into our agentic workflows at Cloudflare. For instance, inside groups need a classifier mannequin to have the ability to consider Trust & Safety submissions, assist us triage Cloudflare Support requests, and even to be built-in to our Bot merchandise to resolve if a crawler is an effective bot or dangerous bot.
These use instances are extremely particular and we’ve had a few years of labelled selections that we may use to coach a particular classifier. When you fine-tune a mannequin, you might surrender some basic objective efficiency in alternate for increased accuracy in a particular area.. Because Cloudflare has greater than 15 years of community knowledge throughout totally different domains, we will fine-tune a mannequin to suit these particular use instances which is extra correct and sooner than our generic Clef mannequin. We’re working with inside groups already to determine how we will post-train Clef to create highly effective ML fashions that enhance our influence and enhance workflows throughout Cloudflare. These inside groups and use instances are the subsequent remit of our new FDE fine-tuning workforce and foundation for our reinforcement studying (RL) product.
Our new RL service
We are providing a service to assist clients fine-tune Clef to go well with their workloads with our hands-on FDE workforce. From that, we’ll study from our hands-on experiences to construct a self-serve platform that clients can use to seize knowledge, fine-tune, and redeploy the mannequin, all on Cloudflare.
This has really been a very long time coming — we’ve been constructing our AI platform to have the proper primitives the place we could possibly be constructing a customized RL product. The curiosity in Jev exhibits the necessity for a quick, small, particular, classifier mannequin, and we selected this to be our area of interest to begin experimenting with RL environments.
To do that, we leverage the primitives that we have already got constructed on our Cloudflare platform:
- Cloudflare AI Gateway – cross all of your AI visitors by AI Gateway and robotically create a dataset of requests on your use case
- Cloudflare Workers AI – generate rollouts towards the bottom Clef mannequin
- Cloudflare Containers – RL sandbox for scoring and replaying agent actions
- [NEW] Trainer – replace weights of fine-tuned Clef mannequin
- Cloudflare Workers AI + BYO Model – redeploy the fine-tuned mannequin on Workers AI
This combines a number of work-in-progress items of the AI Platform that we’ve been engaged on, together with AI Gateway that captures your AI visitors so you possibly can leverage your personal request/response knowledge, Containers for RL Sandboxes, and Workers AI’s Bring Your Own Model (Cog) work that has been progressing since our acquisition of Replicate.

Try it out at the moment
We’re excited to launch our first Cloudflare-trained ML mannequin from the Workers AI workforce at the moment. We’re nonetheless early right here and have much more enhancements in retailer, however it’s a fantastic first showcase of the laborious work we’ve been doing on the AI Platform workforce. We consider that Clef has the power to disrupt the best way we use brokers, which inserts naturally into Cloudflare’s mission of being the agent cloud.
If you may have particular use instances and are already clients of those merchandise — we’d love to chat with you and be design partners as we experiment in this space.
Try out the Clef fashions hosted on Workers AI, obtain the weights on Hugging Face for those who’d wish to probe for your self, and attain out you probably have fine-tuning use instances you’d like us to assist with.
Our ML workforce has been rising in influence, from mannequin optimizations to mannequin coaching analysis. If you’re fascinated about becoming a member of our mission, check out our open roles.
