Cloudflare Releases Clef and Clef-flash: Open-Weight Decision Models That Return Typed Probabilities Instea…
What it does
Cloudflare launched Clef and Clef-flash, two open-weight decision models designed to return typed probabilities instead of generating text. Clef has 27 billion parameters, and Clef-flash is a smaller, 9 billion parameter variant. Both models accept image inputs and are compatible with the Jev-API standard. They run on Cloudflare’s Workers AI platform, delivering median latencies of 209.3 milliseconds for Clef and just 38.8 milliseconds for Clef-flash.
Why it matters
These models shift the AI output form from text generation to returning typed probabilities, which is more structured and machine-readable. This approach can improve decision-making processes that require precise probability assessments rather than open-ended text outputs. The combination of image input support and fast response times makes Clef and Clef-flash practical for real-time, automated decision workflows. Cloudflare’s choice to open-weight these models means developers gain more control and flexibility over how they are used, removing reliance on proprietary APIs or locked-down cloud offerings.
Who it is for
Builders and operators working on decision automation, machine learning inference, or AI-powered pipelines should take note. The Jev-API compatibility ensures easy integration into existing systems. Clef-flash’s low latency makes it suitable for applications needing rapid probabilistic outputs, like real-time content moderation, recommendation systems, or automated image analysis. The more substantial Clef model suits scenarios demanding higher accuracy and complexity at a slightly longer latency.
The catch
These are decision models, not traditional text generation models, so they won’t produce natural language or conversational outputs. Implementing them requires rethinking workflows around probability-based decisions instead of text results. The models’ accuracy and performance in practical tasks remain to be independently evaluated, and running them effectively depends on Cloudflare’s Workers AI environment compatibility and cost structure.
What to watch next
Look for early user feedback on how Clef and Clef-flash perform in production, especially regarding latency, accuracy, and ease of integration. Watch if other providers respond by releasing similar probability-returning models or optimize for typed outputs instead of text. Cloudflare’s investment in AI inference on the edge might pressure competitors to improve speed and API flexibility. Finally, monitor how builders incorporate image input with probabilistic outputs to enable new classes of automation.
AI Quick Briefs Editorial Desk