Live
A railway switch junction where tracks split in two directions
AI & ML

The New AI Models Don’t Talk. They Decide.

Ask Cloudflare’s newest model whether a website is a phishing page and it won’t write you a paragraph. It will return numbers. In Cloudflare’s own example, the model “might classify a domain with a 95% chance it is a fashion website, 85% ecommerce, <1% phishing." That's the whole output.

Cloudflare graphic announcing its Clef decision models
Cloudflare launched its Clef decision models on Oct. 1. Image: Cloudflare

On Oct. 1, three separate teams shipped models built to work this way. Cloudflare released Clef and Clef-flash. Perplexity put out a Decisions API. Strands Labs, which builds Amazon’s open-source agent framework, released Strands Decider 2B. All three call these “decision models,” and all three measured themselves against the same yardstick: a model called Jev that most people outside agent engineering had never heard of a month ago.

The model that started it

Jev comes from a company called TypeSafe. OpenRouter’s listing describes it as “a structured decision model from TypeSafe, and the first of its System One models.” The name is a nod to Daniel Kahneman, as developer Sébastien Dubois explains: fast, intuitive System 1, with ordinary chatbots playing the slow, deliberate System 2. Dubois dates the launch to Sept. 15; OpenRouter lists Sept. 18, likely the date it reached that platform.

The idea is narrow on purpose. You give the model some input and a bounded question, such as yes or no, pick one of these options, or score this on a scale, and it returns probabilities instead of prose. Dubois’s description is the best one going: “a semantic `if` you can afford to call thousands of times.” OpenRouter lists Jev at US$0.042 (CA$0.058) per million input tokens, with output free.

Cloudflare is candid about the debt. Its launch post opens by crediting Jev with introducing “a new decision model concept into the world of AI — a model that produces bounded structured outputs cheaply, quickly and consistently.”

Three answers in one day

Cloudflare’s Clef is two models hosted on Workers AI and open-sourced on Hugging Face under an Apache 2.0 licence. Cloudflare says it froze Qwen3.8-27B for Clef and Qwen3.5-9B for Clef-flash and trained a routing head on top. The speed comes from skipping text entirely: “The decision step is non-autoregressive, so there’s no intermediate text to generate token by token,” the company writes. Across 43 benchmarks it reports a median latency of 209.3 milliseconds for Clef against 524.1 for Jev, and 38.8 for Clef-flash. It claims a 64k context window, double Jev’s. It also describes Clef as fully API-compatible with Jev, which tells you who set the standard. The launch post doesn’t give a price.

Perplexity’s Decisions API runs on a model called pplx-decider-v1-27b. Its documentation supports three question types (yes/no, choice and score), up to 255 options per choice question, and text, JSON or images as input. Price: US$0.04 (CA$0.056) per million input tokens, and “Output tokens are free.” The docs are honest about speed, too: about 2 seconds for a few hundred tokens of input, rising to roughly 23 seconds near the 262,144-token limit. Fast is relative.

Strands Decider 2B is the small one. It’s a two-billion-parameter model meant to run on a laptop, and its announcement reports median latency of about 115 milliseconds on an Nvidia RTX 3090 and 153 on an M3 Mac. Strands released the weights along with “all the training data and scripts we used to build the model,” and says it placed third of 33 in its size class on JevBench.

Why build a model that can’t talk

Agents make a lot of small calls. Which tool next? Is this output acceptable? Should this ticket go to a human? Today many of those calls go to a full language model that writes out an answer, which software then has to parse. Strands lists the jobs it has in mind: model routing, tool selection, evaluations and guardrails.

Decision models attack three costs at once. The answer is a probability, so there’s nothing to parse. It’s cheap enough to call constantly. And because every choice comes with a confidence score, you can log it, set a threshold and send anything uncertain to a person. Our read is that the third one matters most. An agent whose every branch is a logged number is far easier to audit after something goes wrong than one that reasoned its way there in prose.

The trade-off is stated plainly by the people selling it. Strands says its model is “significantly worse at solving complex problems than reasoning models,” and that its inability to generate text “makes it unsuited for coding, chatbots, document summarization, and other common LLM tasks.” These are switches, not brains.

One vendor is missing here. The Neuron reported that OpenAI also introduced a Decisions API as part of a new agent stack on Oct. 1. We could not open OpenAI’s own documentation to confirm its specs, so we’ve left it out until we can.

Everything above is vendor-reported, and every benchmark was chosen by the company that published it. The more telling number is the count of launches. Two weeks after a startup named a category, three companies with far bigger distribution shipped into it on the same Thursday, each one promising to answer Jev’s questions faster.

// Contributor, AI Explained
Catherine Crowe

Catherine Crowe covers AI explained for prompt/power: the plain-English guides that break down how the technology works, what the jargon means and what it changes for everyday people. Originally from Canada, she writes from New Zealand.

Latest from prompt/power

  1. How to Read an AI Company’s S-1: The 7 Numbers That MatterOct 5
  2. OpenAI’s Safety Lead Quit Over Culture. California’s AG Was Already InOct 5
  3. When an AI Agent Breaks In, Who Answers for It?Oct 5
  4. Quebec’s First AI Election: ChatGPT Leaned on an AI-Built Voter GuideOct 5
  5. NYC Is About to Put OpenAI, Anthropic, Google and Meta Under OathOct 4

Leave a Reply

Your email address will not be published. Required fields are marked *