System One Comes to the Agent Stack: Jev, Its Uses, and Its Open-Weight Alternatives

Written by

in

System One Comes to the Agent Stack: Jev, Its Uses, and Its Open-Weight Alternatives

TypeSafe AI’s launch of Jev has spawned dozens of open reproductions in a week — and forced a clearer question than the benchmark race: which decisions in an agent loop ever needed a language model?

by Chris DePuy, 650 Group / September 28, 2026

Abstract concept art of a glowing AI decision core radiating parallel decision paths

On September 14, 2026, TypeSafe AI released Jev, the first public model in what the company calls the System One class: a model that never generates text. An application hands it its state plus typed questions — a Choice over a predefined option list, a Score across described levels, a Noul yes-or-no probability — and the model answers every question in parallel from a single forward pass, returning type-safe values with calibrated probabilities. There is no prose to parse, no JSON repair step, and no schema error surface. The training method, which TypeSafe calls Reinforcement Learning for Calibrated Decisions (RLCD), optimizes for epistemically honest probabilities rather than human-preferred prose. Pricing is $0.042 per million input tokens with output tokens free, and the company puts response time at roughly 70 to 500 milliseconds — a hundredfold gain on both speed and cost against a frontier LLM asked to make the same call.

The pedigree matters here: founder Diogo Almeida, who helped build the research behind ChatGPT at OpenAI and is a co-inventor of RLHF, the technique that made language models follow instructions in the first place. After two years in stealth, he has bet that the missing piece of agent automation was not a better writer, but a decider.

TypeSafe AI: A Decision Primitive, Not a Chatbot

The interface is deliberately narrow. Jev cannot write a briefing or explain its reasoning; it picks among options the application already knows. TypeSafe’s own guidance — ask many questions per call because they evaluate in parallel, give it candidates rather than let it invent them, and treat low-confidence answers as a signal to escalate rather than act — reads less like model documentation than like an operations manual for a classifier that is fast enough to sit inside a hot loop.

Browser Use, Vercel, and LangChain: The First Production Patterns

The most-watched deployment so far is browser-use/jev-ultrafast, an open-source agent from the Browser Use team in which Jev selects the next click from an action space rebuilt from the page DOM at every step, and a small language model wakes up only when a field needs typing. The team’s demo completed a Google Flights search in 7.1 seconds at a reported cost of $0.0039.

Vercel runs the same class of model in production. In the default auto mode of its fx coding agent, a safety reviewer inspects every proposed command; Vercel’s own numbers put that classifier step as much as 18x quicker at the p95 mark than the GPT-5.6-Luna-based model it displaced, with accuracy also improved. LangChain open-sourced the equivalent pattern the following day, documented in its engineering write-up: AutoModeMiddleware gates tool calls before execution, and a companion routing middleware sends each request to the least expensive model equal to the job in about a dozen lines, leaving the routing probabilities in agent state for audit.

Two further patterns suggest where the economics bite hardest. fast-jev-compaction, released by community developer Tamara Tran, scores every tool call in a session and drops the irrelevant ones — standing in for the summarization pass that coding agents trigger whenever context approaches its ceiling, as a relevance filter; one reported run compressed a near-1M-token session to 86K tokens in about one second. And in a map-reduce test, Every’s head of evaluations ran 21 questions across 37 documents in a single request — 777 judgments in under 0.7 seconds for roughly a quarter of a cent. Comparable community runs classified 1,018 research papers for eight cents and triaged 500 emails for 3.5 cents.

The Open-Weight Field: Decider, Laya, Kev, and AnyJev

Closed weights did not survive the week. By day four, an independent field of open reproductions had formed, and four entries carry the most documented evidence.

Mapika/decider ships Apache-2.0 models fine-tuned from Qwen 3.5 bases at 2B, 4B, and a 35B mixture-of-experts, and serves a wire format compatible with TypeSafe’s SDKs, so an existing client can repoint at a local endpoint. Its own documentation is candid about the gap: on the Decision Index panel, decider’s 35B trails Jev within 0.03 on language, retrieval, tools, and arts — and by 0.18 on the knowledge area (GPQA, GSM8K, CRUXEval, MMLU).

Convai Innovations’ Laya is a 421M-parameter, multilingual, non-autoregressive decision model under Apache-2.0 that returns calibrated confidence in a roughly 33-millisecond pass — small and fast enough for per-tool-call gating on commodity hardware. Jared Palmer’s Kev, a family from 0.5B to 8B built on Qwen bases with a TypeSafe-compatible API, targets teams that want to train and run the decider themselves.

The most broadly applicable entry is Nokia’s AnyJev, an Apache-2.0 Python library that turns any open causal LLM into a Jev-shaped decider without retraining, by reading the first-token logits over the declared options. It defines three levels: L0 requires no labels and corrects option-position bias by rotation; L1 calibrates probabilities by temperature scaling on 100 to 500 labeled examples; L2 fits a closed-form head on an intermediate hidden state with 100 to 300 labels and can stop the forward pass early. Its constraint is access: it needs logits or hidden states, which confines it to open models.

The Honest Numbers: Benchmarks and Calibration

For now the closed model still tops the boards. JevBench ranks Jev 1.13.0 first overall at 75.4 across intelligence, calibration, speed, and cost, with the open SemIf second at 74.7 and dozens of entrants behind; the Decision Index leaderboard scored 32 systems over 132,422 requests and 37 benchmarks, again with Jev first at 59.5. Several open systems rank within a few points on intelligence specifically.

The recurring caveat is calibration, not speed. A raw softmax across candidate logits on a repurposed language model is a ranking, not a trained probability — high confidence means the model is sure, not that it is right. TypeSafe’s guidance to threshold on confidence and escalate the uncertain band is what separates a classifier that hands back a label from a decision system that hands back permission to use it; teams that skip the labeled-data step and let an untuned decider auto-act are the failure mode the community keeps documenting.

Where This Lands

We read the split as structural: generation to the language model, decisions to the System One model, execution and policy to code. The open-weight field has closed the interface gap within days — the wire format, the typed question types, even the SDKs now have drop-in equivalents — while the gap that remains is measured knowledge and calibration density, which is exactly the kind of gap a closed lab with a head start defends for one training run, not five years. The buyer-side implication is concrete: the decision layer beneath the agent harness is commoditizing at $0.042 per million tokens, and any agent budget that still pays generation rates on yes-or-no questions is carrying recoverable waste. As firms such as 650 Group have been tracking in their AI infrastructure research, the interesting metric for the next quarter will not be which decider tops JevBench, but how many of the billion daily agent decisions migrate out of the generation tax entirely.

More posts

© MarketIntelligenceResearch.com