A week ago, a model that cannot chat was the most talked-about release in AI. TypeSafe's Jev, which returns a typed decision and a probability instead of prose, drew reported funding talks at a valuation of $10 billion or more, a story we covered in TypeSafe's Jev: The $10B Model That Cannot Chat. Between September 29 and October 1, 2026, the category stopped being a one-company story. Perplexity, Cloudflare, the AWS-backed Strands Agents team, Liquid AI, Inception and Fastino each shipped decision models of their own, several of them with open weights. This piece looks at what the decision models wave actually contains, which numbers are verifiable, and what it means for Jev and for anyone building products on top of many models.
Perplexity
QwenWhat a decision model is, and why the output format matters
A conventional language model produces text one token at a time, and software must then parse that text and hope it is well formed. A decision model is asked a typed question about some input, which the vendors call the state, and returns a structured answer. Perplexity's Decisions API documentation describes three shapes: a yes or no question answered with "the probability of yes, from 0 to 1", a multiple choice answered with a probability for every option plus the most likely one, and a rubric score answered with a probability for every level plus the expected score. Liquid's d1 and Inception's Mercury Decide, as described on their OpenRouter listings, return the same family of answers "with a probability taken directly from the model rather than written out as text."
That last detail is the economic hinge. Because nothing is generated, there are no output tokens to bill. Perplexity charges $0.04 per million input tokens and says output tokens are free; Liquid's d1 is listed on OpenRouter at $0.04 per million input tokens and $0.00 for output; Jev 1.13 is listed at $0.042 per million input tokens. For a router or a guardrail that runs on every request, a price that depends only on how much you send changes the arithmetic of putting a model in the hot path.
Why calibration is the real product
A probability is only useful if it means what it says. If a model reports 90% confidence across a thousand decisions, it should be right about nine hundred times. That property, calibration, is what lets an engineer write a rule such as "act automatically above 0.95, escalate to a larger model or a human below it." A badly calibrated model makes that threshold meaningless, however accurate its top choice is.
The launch posts treat calibration as a first-class metric. Cloudflare says Clef was trained with a method it calls Reinforcement Learning for Calibrated Decisions (RLCD) to refine probability calibration. The Strands team reports its 2B model ranks 3rd of 33 in the 2B class on JevBench's public set, scored by Brier score, a standard measure that penalises confident wrong answers. Fastino takes a different route: its GLiDE model first produces a fast probability distribution, then allocates additional reasoning when the leading result is uncertain, spending compute only on the hard cases.

Who shipped what: open weights against hosted APIs
The most important split in the wave is not speed or score but whether the weights are public. The table below uses only figures stated on each launch page or listing; where a number was not published, it says so.
| Model | Maker | Weights | Listed price | Speed figure (source and method differ) |
|---|---|---|---|---|
| Jev 1.13 | TypeSafe | Hosted | $0.042/M input, output free | 0.16 s P50 round trip on OpenRouter, Oct 3 |
| pplx-decider-v1-27b | Perplexity | Open, Apache 2.0 | $0.04/M input, output free | Under 2 s for a few hundred tokens, Perplexity testing |
| Clef and Clef-flash | Cloudflare | Open, Apache 2.0 | Not stated in launch post | 209.3 ms and 38.8 ms median, Cloudflare-run |
| Strands Decider 2B | Strands Agents (AWS-backed) | Open source, GitHub and Hugging Face | Self-hosted | About 115 ms median on an RTX 3090 |
| d1 | Liquid AI | Hosted | $0.04/M input, output free | 0.43 s P50 round trip on OpenRouter, Oct 3 |
| Mercury Decide | Inception | Hosted | Free for early access on OpenRouter | Up to 14 decisions per second, per OpenRouter listing |
| GLiDE | Fastino | Hosted, Fastino API | Not stated in launch post | Not stated |
Perplexity's model card on Hugging Face lists an Apache 2.0 licence and a Qwen3.8-27B base. Cloudflare publishes Clef on Hugging Face under Apache 2.0 and serves it on Workers AI, and says it is "fully Jev-API compatible." The AWS-linked Strands release, written by Marc Brooker, Mike Chambers and Fabio Nonato de Paula, pitches a 2 billion parameter model "suitable for running on a local CPU or GPU." On the hosted side, OpenRouter describes both d1 and Mercury Decide as "served as a System One endpoint," the label TypeSafe uses for Jev, which suggests the API shape itself is becoming a shared convention.
Median latency per decision, as reported by Cloudflare
Milliseconds, median across 43 benchmarks. A vendor-run comparison published with the Clef launch on October 1, 2026; not an independent measurement.
The latency figures are not comparable across rows. Cloudflare's numbers are model latency from its own benchmark run, in which it measured Jev at 524.1 ms median; OpenRouter's P50 is a live round-trip figure that includes network time and changes hour to hour; Strands measured on a single consumer GPU. The safe reading is directional: several entrants claim sub-second, and in some cases sub-100 ms, decisions.
Every benchmark has a different scoreboard
Nearly every launch claims to beat Jev, and nearly every one uses a different yardstick. Perplexity reports 85.71% across an 11-benchmark panel while trailing Jev on JevBench public hard (70.30% against 73.27%). Cloudflare says Clef leads on the Jev Decision Index and beat Jev in 3 of 4 areas of TypeSafe's own suite, while Jev still led on When2Call accuracy. Fastino reports 64.81 against Jev's published 57.91 on the Decision Index 0.2.1 scorer. Liquid says d1 is the first model to outperform Jev on Hugging Face's Decision Index. These claims cannot all be ranked against each other, and all are self-reported. None of the launch posts points to an independent, shared leaderboard that settles the question.
Nearly every launch claims to beat Jev, and nearly every one uses a different yardstick.
On the state of decision model benchmarks
What decision models are actually used for
The Strands authors list the uses they have seen work: "model routing, tool selection, evaluations, guardrails, memory, context management, and policy classification." Each of these is a point inside an agent where a cheap, well-calibrated yes, no or pick-one is more useful than a paragraph. Cloudflare frames the shift more broadly, arguing that with decision models "a human does not necessarily need to be in the loop for agentic decisions anymore." Whether that holds depends on calibration in production, not on launch benchmarks.

The commoditisation question for Jev
TypeSafe's reported valuation rests on the idea that the decision layer is a large, metered market with a clear leader. Within days, two Apache 2.0 models from well-capitalised companies appeared, one explicitly API-compatible with Jev, and hosted rivals priced at or slightly below Jev's listed rate. That is the classic pattern by which a capability turns into a commodity: open weights set a floor on price, and a shared interface lowers switching cost.
There are counterarguments. Jev was first, is the reference model every rival benchmarks against, and showed the lowest live P50 among the three hosted decision models listed on OpenRouter when we checked. A product company can also differentiate on reliability, tooling and fine-tuning services, which Cloudflare is already offering for Clef. The open question is whether the durable value sits in the model, in the benchmark that defines "good", or in the distribution that puts a decision call in front of every request.
What it means for multi-model products
For applications that already use several providers, a crowded decision layer is mostly good news. A shared typed interface means a router or guardrail can be swapped like any other model, and open weights allow it to run close to the application when latency matters. Products that route requests across many models, such as Metir AI, are the natural consumers of this layer, since deciding which model should answer is itself a decision problem. The practical advice that follows from the launches is modest: treat vendor benchmarks as hypotheses, measure calibration on your own traffic, and keep the decision model replaceable.
The takeaway
In under a week, decision models went from one startup's thesis to a crowded category with at least six new entrants, two Apache 2.0 releases, and input prices clustered around four cents per million tokens. The format, a probability instead of prose, now looks settled. What remains unsettled is who captures the value once that format is everywhere.
Sources:
- Decisions API quickstart | Perplexity
- perplexity-ai/pplx-decider-v1-27b | Hugging Face
- Clef decision models | Cloudflare Blog
- Introducing Strands Decider | Strands Agents
- Introducing GLiDE, the first thinking decision model | Fastino
- Liquid AI d1 announcement thread | Thread Reader App
- LiquidAI: D1 | OpenRouter
- Inception: Mercury Decide (free) | OpenRouter
- TypeSafe: Jev 1.13 | OpenRouter
Image credits
Header image: entrance area of the Cloudflare offices at 101 Townsend Street, San Francisco, photographed through the street window in December 2021, by HaeB, via Wikimedia Commons, licensed under CC BY-SA 4.0. In-body: lava lamp wall at the same Cloudflare office, by HaeB, via Wikimedia Commons, licensed under CC BY-SA 4.0. In-body: Amazon data center in Boardman, Oregon, by Visitor7, via Wikimedia Commons, licensed under CC BY-SA 3.0.
