Early testers use Jev to route requests, check risky agent actions, and grade other models’ work. Vercel reports faster and more accurate safety checks. In one email test, Google’s Gemini was slightly more accurate but 10 to 20 times more expensive.
New to this? Read it in simple words
- Early testers use Jev to make quick decisions inside their software. They use it to route requests, check risky agent actions and grade other models’ work.
- Independent results show Jev close to, but behind, bigger models. In one email test, Gemini was slightly more accurate but cost 10 to 20 times more.
- Jev suits many decisions with a fixed set of answers, where speed and cost matter. It is a poor fit for tasks that need explanations, precise numbers or new text.
- AI agent
- An AI that takes steps on its own to finish a task, such as booking or coding.
- Routing
- Sending each request to the right place, person or tool.
Small decisions, made many times
TypeSafe describes Jev’s answers as “smart if-statements”. Code asks Jev a question, gets a typed answer, and follows it. The company lists classifying, routing, scoring, and extracting as core jobs.
Almeida also expects developers to use Jev as a cheap guard for AI agents. It could watch what an agent does and help stop attempts to trick it into unsafe behaviour.
Speed allows real-time uses too. In one TypeSafe demo, Jev played the video game Doom by making 10 decisions a second, for about $7 an hour. The company admits that a normal game bot without AI could play better.
Early tests show large savings, with accuracy close to, but below, the best language models.
What early testers found
Vercel replaced a safety classifier built on OpenAI’s Luna model with Jev, according to TechCrunch. The classifier reviews commands before they run. With Jev, results came five to 18 times faster and were more accurate.
Bryo AI’s chief technology officer, Nikhil Mudholkar, tested Jev against Google’s Gemini on business emails. Gemini was slightly more accurate, but it cost 10 to 20 times more. He said Jev was the only model that gave back “a real probability”.
At the publication Every, a test ran 777 judgments in under 0.7 seconds for about a quarter of a cent. Jev caught six of seven planted defects, while Claude Fable 5.1 caught all seven.
An open-source tool called pi-warden uses Jev to check each action of a coding agent before it runs. Over 17,000 recorded actions, it stopped 42, and about 88% of those stops were correct.
Where Jev fits, and where it does not
Independent results show Jev close to, but behind, bigger models. On the Jevals site, it scored 69.0 on 300 yes-or-no medical questions, against 73.0 for Gemini 3.8 Flash. The site calls this a statistical tie at 1/28 of the price.
On a banking test that sorts customer requests into categories, Jev scored 67.8 against 74.1 for Gemini.
Cheap language models are rivals too. Good Start Labs puts the cost of a million graded answers at $160 for Jev, $260 for DeepSeek V4.1 Flash, and $33,000 for Claude Fable 5.1.
Jev is a poor fit for tasks that need explanations, precise numbers, or new text. “It delegates the hallucination problem a little bit to the user,” said Armin Ronacher, chief technology officer of Earendil.
Sources
Every fact in this story comes from the sources below. Open them to check our work.
- 1
- 2Research · September 18, 2026A new kind of AI model from a ChatGPT inventor is thrilling developers TechCrunch
- 3Research · September 21, 2026What Is Jev? Inside TypeSafe’s Decision-Only AI Model and Its Developer Use Cases Firecrawl
- 4
- 5
We read TypeSafe’s launch post, then compared it with early tests reported by TechCrunch and Firecrawl and with independent results from Jevals and Good Start Labs. Test sizes differ, and several results come from early users rather than controlled studies.