Cloudflare's Clef brings structured decisions to open models
The new open-weight vision-language family outputs typed, structured results and arrives alongside a reinforcement-learning fine-tuning platform.
Cloudflare has introduced Clef, a family of open-weight decision models designed to take image-and-text input and return structured, typed output rather than freeform prose. The release is paired with a new reinforcement-learning fine-tuning platform, according to Cloudflare's announcement.
The pitch is practical: most vision-language models are built to chat, but many real deployments need a model to make a decision and emit a predictable result that downstream software can consume. Clef is framed as a "decision model" aimed at exactly that gap, producing schema-friendly outputs that fit cleanly into automated pipelines.
Why it matters
Typed, structured output is increasingly where the value is for production AI. Teams routinely wrap chat models in parsing and validation layers to coerce reliable results; a model built to return structured decisions natively reduces that brittleness.
- Open weights published under Cloudflare's clef repository, lowering the barrier to self-hosting and inspection.
- Multimodal input combining images and text, oriented toward vision-language tasks.
- An RL fine-tuning platform, giving teams a path to adapt the models to their own decision criteria.
The combination is notable coming from Cloudflare, whose infrastructure footprint positions it to run such models close to where developers already deploy. For teams that need dependable, machine-readable results from multimodal inputs, Clef is worth a look.
Sources
- Visit
Clef: Open-weight decision models, and new RL fine-tuning platform
Hacker News
More in Vision-Language
Liquid AI's LFM2.5-VL-DSpark targets faster VLM inference
The new vision-language model from Liquid AI is tuned for accelerated inference, extending the company's LFM2 line into multimodal territory.
Xiaomi distills MiMo V2.6 into a 9B model
The new MiMo-V2.6-Distill-Qwen-9B targets agentic workloads, coding, and tool use in a size that fits on modest hardware.
Apple's LensVLM-9B targets long-context vision tasks
A 9-billion-parameter vision-language model built on Qwen3.5-9B leans on visual-text compression to stretch its usable context.
0 comments
No comments yet. Be the first to weigh in.