Intermediate level10 min readTo learn at another level, choose it before you start the course.
After this lesson you canExplain how a decision model differs from a model that writes text, read its probabilities with care, and decide when it may act without a person.
The short answerA decision model does not write text: you give it a question and the possible answers, and it gives each answer a probability. That makes it fast and cheap for repeated choices, such as sorting messages.
In simple words
- You set the question and every allowed answer; the model picks from your list.
- It gives a probability for each answer, not a written explanation.
- Jev, from TypeSafe AI, is a paid service; Amazon’s and Cloudflare’s rivals are free to download.
- A high probability is only useful if tests show it matches how often the model is right.
The short answerA decision model scores a fixed set of answers you define and returns a probability for each, instead of generating text. It suits high-volume routing, checks and grading, if its probabilities are calibrated on your own data and unsure cases go to a person or a larger model.
In simple words
- Jev offers three question types: a choice among up to 255 options, a score on a scale, and the chance that something is true.
- Jev costs $0.042 per million input tokens, and output is free; TypeSafe has not published its design or size.
- Strands Decider 2B and Clef are open-weight rivals, built on Qwen models with a scoring part added.
- Outside tests are mixed: on one site, Jev scored a few points below Gemini 3.8 Flash on yes-or-no questions.
The short answerDecision models expose classification through a typed interface: they score options the developer defines and return a probability distribution, never free text. Their value depends on calibration and on the routing policy around them, and Jev’s architecture is unpublished, so evaluate them on your own data before letting them act.
Key points
- Strands Decider 2B: Qwen3.5-2B with a pointer head of just over one million parameters.
- Clef: a small head and adapters on Qwen base models, trained with a Brier loss and a reinforcement learning step.
- Grade probabilities with the Brier score, where 0 is perfect, and check calibration on its own too.
- Structured outputs constrain the shape of a large model’s answer, not its truth.
1 Where to set the threshold
- I was charged twice for one orderBilling 94%Routed automatically
- Where is my parcel?Delivery 91%Routed automatically
- The app logs me out every hourTechnical 83%Routed automatically
- Can I change the name on my account?Account 71%A person checks
- Refund the broken one, keep the restBilling 58%A person checks
- Your last update deleted my files!Technical 46%A person checks
3 of 6 tickets are routed automatically; 3 go to a person. If the percentages are honest, a higher line means fewer mistakes but more work for people.
2 A writing model and a decision model
| Question | Writing model | Decision model |
|---|---|---|
| You get back | Text, written token by token | Your options, each with a probability |
| You define | A prompt; the answer can be anything | The question and every allowed answer |
| Good for | Writing, explaining, open questions | Routing, checks, ratings, yes or no |
- JevTypeSafe AI · paid service
- Strands Decider 2BAmazon Web Services · open weights
- Clef and Clef-flashCloudflare · open weights
Words to know
- Decision model
- A model that picks from answers you define and gives each a probability, instead of writing text.
- Calibration
- How well the stated confidence matches how often the model is right.
- Threshold
- The confidence line above which the model may act alone; below it, a person checks.
Choosing instead of writing
A chat model writes an answer word by word, and the answer can be anything. A decision model works more like a multiple-choice test. You give it some text, a question and the allowed answers, and it picks one.
TypeSafe AI launched Jev in September 2026 and calls it a “System One” model, after the fast, intuitive thinking that the psychologist Daniel Kahneman described. Jev can choose from up to 255 options, rate something on a scale, or give the chance that a statement is true.
What is inside is not public
TypeSafe says Jev uses a new design that returns all the answers to one request together. It has not published the model’s size, how it was built or a technical report, so nobody outside the company can check how it works inside.
Its rivals are more open. Amazon’s Strands Decider 2B and Cloudflare’s Clef are free to download. Amazon built its model on an existing small language model and swapped the part that writes text for a part that scores each option.
Can you trust the percentage?
A decision model might say “Billing: 94%”. That number is only useful if it is honest: of all the times it says about 90%, it should be right about 9 times in 10. Experts call this calibration.
Jev also gives a confidence: how sure it says it is. TypeSafe itself warns that high confidence does not mean the answer is right. Its documents also list weak spots, such as counting, dates and a habit of picking the first option too often.
So test it on your own examples first. Then set a threshold: let it act alone only above the confidence where your tests show it is usually right, and send the rest to a person. The diagram above lets you try this.
Try it yourself
Ask a chat app to route an email to Billing, Support or Sales, to pick one, and to give a probability for each. The email: “I was charged twice, and now I can’t log in.” Then reorder the options and ask again in a new chat. Did anything change? A chat app’s percentages are only words it writes.
Check yourself
What does a decision model such as Jev give back?
Jev says it is 90% sure. What do you need before you trust that number?
Where does a decision model fit best?
Words to know
- Decision model
- A model that picks from answers you define and gives each a probability, instead of writing text.
- Calibration
- How well the stated confidence matches how often the model is right.
- Threshold
- The confidence line above which the model may act alone; below it, a person checks.
An old idea, a new product
Picking a label is classification, one of the oldest jobs in machine learning. BERT (2018) showed that a pre-trained language model can learn it with just one extra output layer.
Decision models package this for AI agents and apps. You send text and questions of a set type, and the model returns probabilities, never free text. Jev takes up to 64,000 tokens per request and charges $0.042 per million input tokens, with output free.
Three models, three approaches
TypeSafe says Jev uses a new architecture and a “parallel sampler” that returns all answers to one request together. It has published no size, base model, technical report or calibration metric.
A model’s head is its last part, which turns its numbers into an output. Amazon’s Strands Decider 2B is built on Qwen3.5-2B, with the text-writing head swapped for a small “pointer head” that scores each option.
Cloudflare’s Clef (27 billion parameters) and Clef-flash (9 billion) start from Qwen base models. They add a small trained head and adapters, small add-on layers, and also read images.
Both rivals publish their weights under the Apache 2.0 licence. Their speed and accuracy figures are the makers’ own.
Calibration and thresholds
Calibration means the stated confidence matches reality: of 100 answers given at 0.8, about 80 should be right. Guo and others (2017) found that modern deep networks tend to be overconfident, and that a simple fix, temperature scaling, repairs much of it.
TypeSafe’s documents say Jev’s confidence shows how spread out its probabilities are, not whether the answer is correct. Use the confidence, once tested on your own data, to set a threshold: act alone above it, and send the rest to a person or a larger model. To set it, run a few hundred of your own labelled examples and find the confidence above which the model was right often enough.
What outside tests show
On Jevals, a site that says it is unaffiliated, Jev scored 69.0 on yes-or-no medical questions, against 73.0 for Gemini 3.8 Flash. A September 2026 paper found Jev within three points of GPT-6 when the answer is in the text, but behind on maths, code and logic.
The same paper found that accepting Jev’s confident verdicts and passing the rest to a bigger model matched or slightly beat GPT-6 as a judge, at about 41% of the cost. That paper is a preprint, not yet reviewed by other experts. We found no independent test that compares Strands Decider or Clef directly with Jev.
Try it yourself
Ask a chat app to route an email to Billing, Support or Sales, to pick one, and to give a probability for each. The email: “I was charged twice, and now I can’t log in.” Then reorder the options and ask again in a new chat. Did anything change? A chat app’s percentages are only words it writes.
Check yourself
A team routes support tickets with a decision model. What is the safest way to use its probabilities?
What has TypeSafe published about how Jev works inside?
Cloudflare says a Clef model scores highest on 7 of 10 decision tests. How should you read that?
Interfaces and internals
Jev exposes three primitives: Choice, one of up to 255 options with a probability for each plus a confidence; Score, 2 to 10 ordered levels; and Noul, the chance of yes. TypeSafe claims a new architecture, a parallel sampler and a training method it calls RLCD, but publishes no size, base model or report.
The open rivals document their designs. Strands Decider 2B replaces Qwen3.5-2B’s language-model head with a pointer head of just over one million parameters that scores each option.
Clef and Clef-flash train a small head and adapters on Qwen3.8-27B and Qwen3.5-9B, with a Brier loss and a reinforcement learning step that Cloudflare also calls RLCD. Neither source makes clear whether it is the same method as TypeSafe’s.
Calibration and scoring rules
Guo and others (2017) showed that modern deep networks tend to be overconfident, and that temperature scaling fixes much of this without changing accuracy. The Brier score (1950) grades probability forecasts by squared error, so 0 is perfect. It mixes calibration with accuracy, so check calibration on its own too, for example with a reliability chart.
AWS reports its own JevBench run for Strands Decider: 167 of 231 right, or 72.3%, with a calibration error of 0.050. TypeSafe reports no calibration metric for Jev, and its documents say that confidence measures spread, not correctness.
Generative models can decide too
A large model can be made to choose with structured outputs, which force the answer into a JSON schema, including fixed lists. Constrained decoding allows, at each step, only tokens that keep the output valid.
That fixes the shape, not the truth: OpenAI warns that structured results can still contain mistakes. Token probabilities can serve as confidence, but they are not correctness either.
Evidence and cascades
Cloudflare’s own tables show the spread. On BANKING77, a test of sorting banking questions into 77 types, Clef scores 94.2 and Jev 79.7; on GPQA Diamond, a reasoning test, Jev 78.3 and Clef 48.0. Cloudflare also claims a median of 524.1 ms for Jev against 38.8 ms for Clef-flash, without describing its method.
Li and others (2026) found Jev within three points of GPT-6 when the answer is in the text, but behind on maths, code and logic. In this preprint, accepting confident verdicts and escalating the rest matched or slightly beat GPT-6 as a judge, at about 41% of the cost. RouteLLM (2024) cut costs by over 2 times in some tests by routing between a strong and a weak model.
TypeSafe lists Jev’s own weak spots: literal reading, counting and dates, double negatives, long irrelevant text, injected instructions and a habit of picking the first option too often.
Try it yourself
Ask a chat app to route an email to Billing, Support or Sales, to pick one, and to give a probability for each. The email: “I was charged twice, and now I can’t log in.” Then reorder the options and ask again in a new chat. Did anything change? A chat app’s percentages are only words it writes.
Check yourself
Of 200 tickets a model labelled at 0.9 confidence, 150 were right. What does that show?
A team uses structured outputs to force a large model to answer only “approve” or “reject”. What does that guarantee?
TypeSafe says Jev picks the first option too often. What should an evaluation do?
Sources
- Introducing System One models and Jev (TypeSafe AI, September 2026)
- TypeSafe emerges from stealth (DCVC, September 2026)
- Jev models, limits and prices (TypeSafe documentation)
- Choice, Score and Noul question types (TypeSafe documentation)
- Confidence is not correctness: the Score type (TypeSafe documentation)
- The Noul type: the chance that a statement is true (TypeSafe documentation)
- Jev’s known weak spots (TypeSafe documentation)
- Introducing Strands Decider (Strands Agents, AWS, October 2026)
- Strands Decider 2B model card (Hugging Face)
- Clef decision models (Cloudflare, October 2026)
- Clef on Workers AI (Cloudflare documentation)
- Clef model card (Cloudflare, Hugging Face)
- Clef on Workers AI: decision benchmarks (Cloudflare changelog, October 2026)
- Jevals: outside tests of Jev
- JEV-as-a-Judge (Li and others, 2026, arXiv)
- BERT: Pre-training of Deep Bidirectional Transformers (Devlin and others, 2018, arXiv)
- On Calibration of Modern Neural Networks (Guo and others, 2017, arXiv)
- Verification of Forecasts Expressed in Terms of Probability (Brier, 1950, Monthly Weather Review)
- Efficient Guided Generation for Large Language Models (Willard and Louf, 2023, arXiv)
- Structured outputs (OpenAI API documentation)
- Using log probabilities (OpenAI Cookbook)
- RouteLLM: Learning to Route LLMs with Preference Data (Ong and others, 2024, arXiv)
This lesson was generated by AI systems under the editorial rules of Silicon AI News and checked against the sources it lists. The live parts come from our checked stories, trackers, model comparison and rules checker.