Intermediate level10 min readTo learn at another level, choose it before you start the course.

After this lesson you canExplain how a decision model differs from a model that writes text, read its probabilities with care, and decide when it may act without a person.

The short answerA decision model scores a fixed set of answers you define and returns a probability for each, instead of generating text. It suits high-volume routing, checks and grading, if its probabilities are calibrated on your own data and unsure cases go to a person or a larger model.

In simple words

  • Jev offers three question types: a choice among up to 255 options, a score on a scale, and the chance that something is true.
  • Jev costs $0.042 per million input tokens, and output is free; TypeSafe has not published its design or size.
  • Strands Decider 2B and Clef are open-weight rivals, built on Qwen models with a scoring part added.
  • Outside tests are mixed: on one site, Jev scored a few points below Gemini 3.8 Flash on yes-or-no questions.
Choosing instead of writingIllustration + live data

1 Where to set the threshold

  1. I was charged twice for one orderBilling 94%Routed automatically
  2. Where is my parcel?Delivery 91%Routed automatically
  3. The app logs me out every hourTechnical 83%Routed automatically
  4. Can I change the name on my account?Account 71%A person checks
  5. Refund the broken one, keep the restBilling 58%A person checks
  6. Your last update deleted my files!Technical 46%A person checks

3 of 6 tickets are routed automatically; 3 go to a person. If the percentages are honest, a higher line means fewer mistakes but more work for people.

2 A writing model and a decision model

QuestionWriting modelDecision model
You get backText, written token by tokenYour options, each with a probability
You defineA prompt; the answer can be anythingThe question and every allowed answer
Good forWriting, explaining, open questionsRouting, checks, ratings, yes or no
  1. JevTypeSafe AI · paid service
  2. Strands Decider 2BAmazon Web Services · open weights
  3. Clef and Clef-flashCloudflare · open weights
The tickets and percentages are made up to show the idea. The models come from our model release tracker; their speed and accuracy figures are the makers’ own.

Words to know

Decision model
A model that picks from answers you define and gives each a probability, instead of writing text.
Calibration
How well the stated confidence matches how often the model is right.
Threshold
The confidence line above which the model may act alone; below it, a person checks.

An old idea, a new product

Picking a label is classification, one of the oldest jobs in machine learning. BERT (2018) showed that a pre-trained language model can learn it with just one extra output layer.

Decision models package this for AI agents and apps. You send text and questions of a set type, and the model returns probabilities, never free text. Jev takes up to 64,000 tokens per request and charges $0.042 per million input tokens, with output free.

Three models, three approaches

TypeSafe says Jev uses a new architecture and a “parallel sampler” that returns all answers to one request together. It has published no size, base model, technical report or calibration metric.

A model’s head is its last part, which turns its numbers into an output. Amazon’s Strands Decider 2B is built on Qwen3.5-2B, with the text-writing head swapped for a small “pointer head” that scores each option.

Cloudflare’s Clef (27 billion parameters) and Clef-flash (9 billion) start from Qwen base models. They add a small trained head and adapters, small add-on layers, and also read images.

Both rivals publish their weights under the Apache 2.0 licence. Their speed and accuracy figures are the makers’ own.

Calibration and thresholds

Calibration means the stated confidence matches reality: of 100 answers given at 0.8, about 80 should be right. Guo and others (2017) found that modern deep networks tend to be overconfident, and that a simple fix, temperature scaling, repairs much of it.

TypeSafe’s documents say Jev’s confidence shows how spread out its probabilities are, not whether the answer is correct. Use the confidence, once tested on your own data, to set a threshold: act alone above it, and send the rest to a person or a larger model. To set it, run a few hundred of your own labelled examples and find the confidence above which the model was right often enough.

What outside tests show

On Jevals, a site that says it is unaffiliated, Jev scored 69.0 on yes-or-no medical questions, against 73.0 for Gemini 3.8 Flash. A September 2026 paper found Jev within three points of GPT-6 when the answer is in the text, but behind on maths, code and logic.

The same paper found that accepting Jev’s confident verdicts and passing the rest to a bigger model matched or slightly beat GPT-6 as a judge, at about 41% of the cost. That paper is a preprint, not yet reviewed by other experts. We found no independent test that compares Strands Decider or Clef directly with Jev.

Try it yourself

Ask a chat app to route an email to Billing, Support or Sales, to pick one, and to give a probability for each. The email: “I was charged twice, and now I can’t log in.” Then reorder the options and ask again in a new chat. Did anything change? A chat app’s percentages are only words it writes.

Check yourself

  1. A team routes support tickets with a decision model. What is the safest way to use its probabilities?

  2. What has TypeSafe published about how Jev works inside?

  3. Cloudflare says a Clef model scores highest on 7 of 10 decision tests. How should you read that?

Sources

This lesson was generated by AI systems under the editorial rules of Silicon AI News and checked against the sources it lists. The live parts come from our checked stories, trackers, model comparison and rules checker.