AIAnalysis

AI decision models: AWS Strands Decider vs Cloudflare Clef

AI decision models pick from a fixed list and return confidence scores, not text. How AWS's Strands Decider and Cloudflare's Clef compare, and who needs one.

Source-based. Written from the documents, reporting and reviews linked in the text. Nothing here was tested hands-on by The Ruling Desk. How we work

A Gigabyte GeForce RTX 3090 graphics card with three fans standing on a wooden desk
Photo: PantheraLeo1359531 / Wikimedia Commons, CC BY 4.0

AI decision models are a new kind of model that never writes a sentence: you give one some context and a short list of allowed answers, and it hands back the answer it picked plus a confidence score for every option. On October 1, 2026, Amazon Web Services released an open one called Strands Decider 2B and Cloudflare released two called Clef and Clef-flash, days after OpenAI previewed its own. Here is what these models do in plain words, how the two open releases compare on size, license, where they run and the vendors' own speed claims, and which developers would actually use one.

Key takeaways

  • A decision model chooses, it doesn't write. It answers questions like "which team should get this ticket?" by scoring a fixed set of options, so your code gets a label and a probability instead of a paragraph to parse.
  • Strands Decider 2B is small enough for a laptop. It's a 2-billion-parameter model on Alibaba's Qwen3.5-2B, released under the Apache 2.0 license, and AWS says it answers in a median 115 milliseconds on an RTX 3090 graphics card.
  • Cloudflare's Clef is bigger and hosted. Clef (27B) and Clef-flash (9B) are also Apache 2.0, run on Cloudflare's Workers AI, and Cloudflare claims Clef-flash has a median latency of 38.8 ms. Those are the vendors' numbers, measured on different hardware, so they don't compare directly.
  • Both copy a startup's idea. TypeSafe AI's paid Jev model started the category in September, and Clef is built to accept the same API calls.
  • Most people will never touch one. These are building blocks for developers running AI agents and high-volume automation, not chatbots.

What AI decision models actually do

A regular chatbot model generates text one word at a time. That's great for an essay and wasteful for a yes or no. If an agent only needs to know whether an email is a refund request, a full language model spends time and tokens writing an answer that your code then has to parse, and it can phrase the answer differently each time.

A decision model skips the writing. Cloudflare's Clef announcement describes it as a model that "makes classifications to help agents decide how to act," returning typed answers with probabilities that your code can use to route a ticket, escalate it or hand it to a human. The model reads the whole input in one pass and then scores the valid choices in parallel, instead of producing tokens one by one.

The pitch is speed, consistency and a usable confidence number. If the model is 97% sure, your workflow moves on; if it's 55% sure, you can send the case to a bigger model or a person.

Where the idea came from

The template is Jev, from TypeSafe AI. TechCrunch's Tim Fernholz reports that Jev launched in September 2026 and that the decision-model products that followed are, in effect, clones of it. TypeSafe's site calls Jev its first public "System One Model," built for decisions inside software, and lists its price at $42 per billion input tokens. Jev is a paid, hosted API; the AWS and Cloudflare models are open weights you can download.

How Strands Decider 2B works

Strands Decider is published by Strands Labs, and TechCrunch credits Amazon Distinguished Engineer Marc Brooker with the work. According to the Strands Decider GitHub repository, the team took Qwen3.5-2B-Base (1.9 billion parameters), threw away the part that generates text, and replaced it with a small "pointer head" of about a million parameters that scores each option. A light fine-tune (a rank-16 LoRA adapter) teaches the rest of the model to feed that head.

It handles three question types: yes or no, one choice out of several, and a score on an ordered scale. The repository lists these numbers for the current reference version, v19:

  • Accuracy: 0.723 (167 of 231) on the public set of JevBench, the benchmark named after Jev.
  • Latency: a median of 115 ms and a 95th percentile of 299 ms on an Nvidia RTX 3090; a warm median of 153 ms on an Apple M3 Pro for inputs under 300 tokens.
  • Calibration: answers given with 0.9 confidence or more are right about 95% of the time, AWS says.

You install it with pip install strands-decider and either ask it questions from the command line or run it as a local HTTP server, on Nvidia, Apple silicon or a plain CPU. The weights are on Hugging Face under Apache 2.0, which allows commercial use. SiliconANGLE reports that AWS picked the 2-billion-parameter size on purpose, saying it strikes the right balance, and pitches it for model routing, tool selection, guardrails and policy checks. TechCrunch adds that it briefly ranked first on JevBench among models its size.

How Cloudflare Clef works

Cloudflare's two models are bigger. According to the company's post by Michelle Chen, Clef is built on Qwen3.8-27B and Clef-flash on Qwen3.5-9B, each with a routing head trained alongside rank-256 low-rank adapters. Both are open-sourced under Apache 2.0 and hosted on Workers AI, Cloudflare's service for running models on its own GPUs.

The Clef-flash model card says it accepts text, JSON, images or video, and that its API is "fully compatible with Jev." That matters more than it sounds: a team already calling Jev can, in principle, point the same code at Cloudflare.

Cloudflare's own numbers, measured on its hosted service:

  • Median latency: 209.3 ms for Clef, 38.8 ms for Clef-flash and 524.1 ms for Jev.
  • Benchmarks: on the BANKING77 intent test, a macro-F1 score of 94.20 for Clef, 90.93 for Clef-flash and 79.74 for Jev. Cloudflare also says Clef beat Jev in 3 of 4 areas of TypeSafe's own evaluation suite.

These are Cloudflare's figures, not independent tests. The post doesn't publish a Workers AI price for Clef.

Strands Decider vs Clef, side by side

Strands Decider 2BClef-flashClef
Made byAWS (Strands Labs)CloudflareCloudflare
Size1.9B (Qwen3.5-2B-Base)9B (Qwen3.5-9B)27B (Qwen3.8-27B)
LicenseApache 2.0Apache 2.0Apache 2.0
Where it runsYour laptop or serverWorkers AI, or self-hostedWorkers AI, or self-hosted
InputsTextText, JSON, images, videoNot detailed
Vendor's median latency115 ms on an RTX 309038.8 ms on Workers AI209.3 ms on Workers AI
PriceFree to downloadNot publishedNot published

The latency row is the trap. AWS measured Strands Decider on a single consumer graphics card; Cloudflare measured Clef on its own data center GPUs. Neither company tested the other's model, and JevBench and Cloudflare's benchmark set aren't the same test, so you can't call one faster or more accurate from these numbers.

More AI decision models from OpenAI and Perplexity

AWS and Cloudflare weren't alone. At DevDay on September 29, OpenAI announced a Decisions API on GPT-6 Luna, posted by its developer account as "available in limited preview," with no published price. Our DevDay 2026 roundup covers it alongside the rest of the event, and our GPT-6 Luna explainer covers the model underneath.

Perplexity followed on October 1 with pplx-decider-v1-27b, another Qwen3.8-27B fine-tune whose weights are on Hugging Face under Apache 2.0. Perplexity reports an overall 85.71% against Jev's 84.51% on its own benchmark panel, and its developer account lists the price at $0.04 per million input tokens. So within weeks of Jev's launch, four big companies had a version, and three of the four released open weights.

What it means for you

If you build software that calls an AI model thousands of times a day to sort, route or approve things, decision models are worth a test. The kinds of jobs they suit:

  • Routing: send a support ticket to billing, tech or sales.
  • Agent steps: pick the next tool, or decide whether a task is done.
  • Guardrails: flag a message that breaks a policy before it goes out.
  • Triage with a fallback: let the small model handle the confident cases and escalate the rest.

Which one fits depends on where your code already lives:

  • Strands Decider 2B suits teams that want to keep data on their own machines, work offline, or avoid a per-call bill. A 2B model runs on a laptop, and Apache 2.0 lets you ship it inside a product.
  • Clef or Clef-flash suits teams already on Cloudflare Workers, or anyone using Jev who wants an open alternative with the same API. Clef-flash also takes images and video.
  • Neither is for open-ended answers. If you need a summary, an explanation or extracted free text, you still need a regular language model.

If you just use ChatGPT or Claude, nothing changes for you directly. You may notice it indirectly, if the apps you use get faster and cheaper because a tiny model is making the routine calls behind the scenes. For more on the models behind those apps, see our AI section.

Bottom line

AI decision models are a narrow, useful idea: let a small model pick from a list and say how sure it is, and save the big model for the hard cases. Strands Decider 2B is the one to try if you want something free that runs on your own hardware; Clef is the one to try if you're on Cloudflare or already speak Jev's API. Every speed and accuracy figure so far comes from the companies themselves, measured on their own hardware and benchmarks, so the next thing to watch is independent side-by-side testing, plus prices for Clef on Workers AI and OpenAI's Decisions API.

FAQ

What is a decision model in AI?

It's a model that chooses an answer from a fixed list you provide and returns a confidence score for each option, instead of writing text. Cloudflare describes it as a model that makes classifications to help agents decide how to act. Developers use them for routing, classification and next-step choices in automated workflows.

Is Strands Decider 2B free?

Yes. The code is on GitHub and the weights are on Hugging Face under the Apache 2.0 license, which allows commercial use. You pay only for the hardware you run it on.

Can I run Clef on my own computer?

The weights for both Clef models are published under Apache 2.0, so you can self-host them, but they are 9 and 27 billion parameters and need far more GPU memory than Strands Decider's 2 billion. Cloudflare's main offering is running them on Workers AI.

How is a decision model different from structured output in a normal LLM?

A normal model with structured output still generates its answer token by token and then formats it. A decision model, as AWS and Cloudflare describe theirs, never generates text: it scores the allowed options directly in one pass, which is where the vendors' speed claims come from.

Filed under AI

Newsletter

New articles, in your inbox.

Free. Unsubscribe in one click. Your email is kept by beehiiv, our newsletter service, and used only for this newsletter.