Three Jev Alternates Landed in a Week: Cloudflare Clef, JEV-27B-VL, and the Open Decision-Model Family

Written by

in

by Chris DePuy, 650 Group / October 2, 2026

Three developments inside a week have moved decision models — the System One models that return typed, calibrated probabilities instead of generated prose — from a single hosted API into a crowded open-weights field. Cloudflare shipped two Apache-2.0 decision models, Clef and Clef-flash, post-trained from frozen Qwen backbones. AutoTrust AI shipped JEV-27B-VL, which puts vision behind the same typed-decision surface. And the open family around Jev keeps widening: a trainable Kev line at 0.6B to 8B, a completion-claim guardrail called Canny, and a DiffusionGemma integration that merged into vLLM upstream on September 22. I think the open-weights decision-model market went from one vendor to a competitive field inside a week, and the calibration receipts each vendor published are what tell you who did the work.

Cloudflare's Clef decision models turn a state and a typed question schema into calibrated probabilities
Cloudflare’s Clef. Source: Cloudflare.

Cloudflare Clef: a hosted edge play with open weights underneath

Cloudflare’s announcement describes Clef as a 27B multimodal model that turns a state and a schema of typed questions into decisions, reading the state as text, JSON, images, or video and returning a probability for every allowed option of every question in a single forward pass, with no free-form generation and no output parsing. The API is fully compatible with Jev and SystemOne, so the swap is a base_url change. Two engineering choices stand out. The backbone is frozen — Qwen3.8-27B for Clef, Qwen3.5-9B for Clef-flash — with only a rank-256 low-rank adapter and the joint routing head trained. And post-training pairs label-smoothed cross-entropy on valid schema outputs with Brier loss to refine probability calibration, so the probabilities themselves are shaped, not just the argmax. Cloudflare also claims a 64k context window against Jev’s 32k.

On the Decision Index 0.2.1 suite the hosted evaluation puts Clef ahead of Jev on most rows — BANKING77 macro-F1 of 94.2 against Jev’s 79.7, median latency of 209.3 milliseconds against 524.1 — though When2Call stays Jev’s win at 81.0 against 72.4. On four end-to-end business workflows, Clef leads invoice processing at 64.7 exact actions against Jev’s 61.8 and security incidents at 62.9 against 61.7, while Jev keeps agent trace observability at 71.6. The caveat is that this is the vendor’s own index, run on its own hosting: an operator moving the swap onto its own GPUs should expect to re-run the suite before trusting the numbers, and the open weights on Hugging Face make that possible rather than theoretical.

JEV-27B-VL: vision arrives on the decision surface

AutoTrust’s card is the most detailed receipt of the three. A serve_decide.py wrapper adds a plain-HTTP POST /v1/decide route to a stock vLLM server. Choice questions scale to 256 options with no retraining — zero-shot on CLINC150 with all 150 intents as options, 89.5 percent with intent names alone and 93.8 percent with a one-line description per option. Needle decisions hold at depth: one sentence hidden at a random depth in up to 250,000 tokens of text was found 20 of 20 times at every length tested. The demo reel is decision making on pictures — Super Mario at roughly 150 milliseconds per decision, a Rubik’s Cube at about 46, Tetris at about 122.

Two results do the positioning. On MicroLens-100k short-video recommendation from cover images alone, with no interaction logs, JEV-27B-VL System 1 reaches AUC 0.727 — statistically level with item-based collaborative filtering learned from 59,045 users’ watch histories (0.728), with a higher top-5 hit rate, 59 percent against 49. That is the cold-start argument for feeds: a brand-new video can be ranked from its cover the moment it is uploaded. As an agent judge on Plan-RewardBench it reaches 73.2 percent macro-average across 1,171 trajectory pairs, the top of that paper’s table, ahead of Qwen-Plus at 70.0, DeepSeek-V3.2 at 69.6, and GPT-5 at 68.5. Calibration is quantified rather than asserted: KL divergence of about 0.017 from TypeSafe Jev 1.13’s distributions, yes/no AUROC of 0.995, choice top-1 agreement of 95.8 percent, rating error of 0.098 on a 0-to-5 scale, and an expected calibration error of 0.0009, per the card’s fidelity table.

The family around Jev is the real story

The Kev family — 0.6B, 4B, and 8B — is the self-serve lane: small Qwen3-based models with a LoRA and a pointer head, a drop-in TypeSafe System One API, and published out-of-domain numbers of 79.6 percent for Kev-8B against 85.7 percent for Jev. Kev-4B serves on a 32 GB Mac in bf16 at roughly 300 milliseconds for five questions, about 40 milliseconds on an H100, and trains in 40 minutes on one H100 under an Apache-2.0 license. Canny applies the same judgment surface to coding agents — reading tool outputs, diffs, and test results to score whether an agent’s claim that it is done is actually reliable. And Google’s open DiffusionGemma 26B-A4B (about 25.2B total parameters, 3.8B active) gained a structured-generation mode in vLLM that fills bounded answer slots with probabilities in parallel; the developer reports about 120 milliseconds for a single decision request and roughly 162 decisions per second at concurrency 32 on one DGX Spark, with an NVFP4 build near 18.9 GB fitting a 24 GB GPU. That patch merged upstream on September 22, so this is no longer a side patch but mainline capability.

The throughline is that the decision layer has become its own market segment with choices at every size — a 9B flash model at the edge on Workers AI, 27B vision models locally, a trainable family for teams that want to own the weights — and that the credible vendors are the ones publishing their calibration receipts rather than only their accuracy. As firms such as 650 Group have been tracking in their AI infrastructure research, value keeps migrating into the machinery around the model, and this week that machinery learned to show its probabilities.

More posts

© MarketIntelligenceResearch.com