In brief

Jev reads a piece of text and answers questions whose possible answers the developer defines. It can pick an option, rate something on a scale, or give the chance that a statement is true. TypeSafe has not published the model’s design or size.

New to this? Read it in simple words
  • A program sends Jev some text and a list of questions. Jev answers them with probabilities, not with sentences.
  • It can pick an option, rate something, or give the chance that a statement is true. TypeSafe says a full answer takes 70 to 500 milliseconds.
  • This way of working makes Jev fast and cheap. But because it never writes sentences, it cannot explain its answers.
  • TypeSafe has not published the model’s design, size or calibration results. Teams should check how often Jev is right when it sounds sure.
Words to know
Probability
A number that shows how likely something is.
Millisecond
One thousandth of a second.
Calibration
How well an AI’s confidence matches how often it is really right.

Three kinds of questions, and no free text

Developers send Jev a “state”, which is text such as a message, a document, or JSON data. They also send named questions. Jev answers all of them in one call.

A Choice question picks one option from a list of up to 255. Jev returns the chosen option, a probability for every option, and a confidence number.

A Score question rates something against ordered levels that the developer describes. A Noul question asks yes or no, and it returns the chance that the answer is yes.

Jev never writes sentences, so it cannot explain why it chose an answer. Developers must build that part of the work in their own code.

Sources14

DECISION PATH 02
Text and questions go in. Typed answers with probabilities come out.

TypeSafe says answers take 70 to 500 milliseconds. The model’s design and size are not public.

Fast answers from a design that stays secret

TypeSafe’s launch post says a full answer takes 70 to 500 milliseconds. Its documentation allows 64,000 tokens per request, but only 32,000 for the text plus the longest question.

The company says Jev uses a new model architecture and a parallel sampler. It also describes a training method called Reinforcement Learning for Calibrated Decisions, or RLCD.

Almeida told the Latent Space podcast that all of its training data is synthetic. TypeSafe has not published the model’s size, its base model, or a technical report.

TechCrunch reports that outside observers suspect Jev is built on an open-weight language model. The company has not confirmed this.

Sources4256

Calibrated does not mean correct

TypeSafe says Jev “can’t hallucinate”, but the claim is narrow. It means every answer fits the format the developer asked for. It does not mean every answer is right.

The documentation lists weak spots. Jev can be quite literal, struggles with numeric precision, and has trouble with indirect questions. Unrelated text in the input can lower its accuracy, and injected instructions can move its answers.

Outside tests show a gap to the largest models. In 6,003 rubric checks by Good Start Labs, Jev gave the same verdict as Claude Fable 5.1 in 91.5% of cases.

The large language models in that test agreed with each other 88% to 95% of the time. Jev agreed with each of them 86% to 92% of the time.

Sources437

Sources

Every fact in this story comes from the sources below. Open them to check our work.

  1. 1
    Primary source · September 2026Choice TypeSafe AI
  2. 2
    Primary source · September 2026Models TypeSafe AI
  3. 3
    Primary source · September 17, 2026Jev 1.13 jaggedness TypeSafe AI
  4. 4
    Primary source · September 15, 2026Introducing System One Models & Jev TypeSafe AI
  5. 5
  6. 6
  7. 7
    Research · September 15, 2026Verification is the bottleneck Good Start Labs
How we checked this story

We read TypeSafe’s documentation on question types, models, and known weaknesses, and its launch post. We compared these claims with TechCrunch, the Latent Space interview, and an independent test by Good Start Labs. Speed and training details are company claims that no outside group has checked.