lyr.ai

New tech, explained

The question mattered more than the model

Jev is a new kind of AI model from a startup called TypeSafe. It doesn’t write text. You give it some input and a few typed questions (pick one of these options, score this, is this true?) and it returns answers with probabilities.

It launched with big numbers: “193.6x faster”, “can’t hallucinate”. I skipped those and read the independent tests instead. Two of them, by different people, point the same way:

How you ask mattered more than which model answered.

Slope chart. Asked one question, 'is this phishing?', Jev scores 62.6% and Haiku, a small LLM, 81.3%. Given five narrow questions with weights fitted on 1,000 labelled emails, Jev scores 95.0% and Haiku 93.2%, statistically tied (p = 0.063). A dashed line shows a regex at 91.8%. Jev is about 27 times cheaper and 5 times faster here, at comparable quality.
Figure 1. Splitting the question moved both models more than switching models did. On this synthetic dataset a regex already reaches 91.8%, and the five questions were written after reading how the dataset was built.

What Jev is

Jev’s documentation describes three kinds of question: a choice from a list, a score on a rubric, and a true/false question, which returns a probability. Several questions can go in one request, and each is answered “in parallel and in isolation”. Input costs $0.042 per million tokens, and output is free.

The docs also say how it’s meant to be used: “Atomic questions, composed in code.” Keep each question narrow, and combine the answers in your own code.

That advice turns out to be the whole story.

Where would I use it?

Vercel, which serves Jev through its AI Gateway, draws the line in one sentence: “Choose Jev for bounded decisions, code for fixed rules, and a generative model for prose.” In a real agent or app, that looks like this:

Decision point Jev? Why
Route a request or ticket to the right team or agent Yes A choice from a known list
Review a proposed tool call: run it, or pause for approval Yes, as one input A yes/no with a probability; the permission itself stays in code
Score a generated answer against stated requirements Yes A bounded score
Write the reply to the user No “Jev does not generate prose”; use an LLM
Check a fixed rule, like balance > 100 No Ordinary code is exact and free
Decide an open-ended plan or strategy Probably not My judgment: it needs reasoning that doesn’t reduce to a few fixed options

The first three rows are uses Vercel lists (alongside prioritizing tickets, categorizing documents, moderation and choosing which model answers). Vercel also says the guardrail doesn’t replace permissions: “Enforce access rules and required approvals in application code before executing a tool.”

The short version: use it at bounded decision points, not everywhere you’d use an LLM. The tests below show how well it holds up there.

Test 1: one question or five

An independent benchmark (jev-phishing-bench) ran Jev and Claude Haiku 4.5 on 2,000 synthetic emails, half of them phishing.

Asked the single question “is this phishing?”, Jev got 62.6% right and Haiku 81.3%. On that framing, the ordinary LLM clearly wins.

Then the author asked both models five narrow questions (for example, whether the sender looks generic), and fitted a small logistic regression on 1,000 labelled emails to combine the answers. Jev reached 95.0% and Haiku 93.2%. The benchmark calls that “statistically tied” (p = 0.063), and Haiku’s AUROC was slightly higher.

The biggest change wasn’t which model answered. It was how the task was decomposed.

What Jev kept was cost and speed. The benchmark’s own summary: “about 27 times cheaper and 5 times faster than Haiku for signals of comparable quality”. That’s the real case for a model like this: it makes asking many narrow questions cheap.

Two caveats the benchmark states itself, and which matter:

Test 2: “which one?” or “is it this one?”

A second independent study (jev-does-not-play-dice) asked Jev about events nobody can predict, like a fair die roll or a coin flip.

Asked as a choice, “which face will it show?”, Jev was badly overconfident. Over 400 die rolls its average confidence was 82.9% and it was right 19.0% of the time, about what guessing gives. On coin flips it said 92% and was right 52%.

Asked as a yes/no, “is it this face?”, its answers were much closer to the truth: an average of 19.2% for an event with a true probability of 16.7%. It isn’t perfect. For rarer events it drifts high, saying about 15% for something that happens 5% of the time.

Dot plot. Asked which face a die will show, the model said 82.9% and was right 19% of the time; asked heads or tails, it said 92% and was right 52% of the time. Asked whether a die will show a given face, it said 19.2% against a true 16.7%; for a 1-in-20 event it said 15% against a true 5%.
Figure 2. Same uncertainty, asked two ways. The confidence you get back depends on the form of the question.

So the same model, facing the same uncertainty, gave very different confidence depending on how the question was framed.

On a real task the picture can be much better: a separate test on 60 hand-labelled agent tool calls got 91.7% right, and every wrong answer came with confidence below 1. Confidence depends on the task and on the question. It has to be checked, not trusted.

“Can’t hallucinate” means “can’t leave the schema”

TypeSafe says Jev “can’t hallucinate”. Its own launch post qualifies this: “Our number is not empirical. Schema matching is guaranteed.”

That’s the precise meaning. Jev can only answer with the options you gave it. It can’t invent a category or return malformed output. But it can still pick the wrong option, confidently, as the die shows.

What I’d take from this

If you use a decision model like Jev:

What I don’t know


Sources: TypeSafe’s launch post and documentation; Vercel’s guides “When should you use Jev instead of a chat model?”, “7 practical Jev use cases” and “Where does Jev fit in an AI agent loop?”; jev-phishing-bench (github.com/anisselbd/jev-phishing-bench); jev-does-not-play-dice (github.com/KantaHayashiAI/jev-does-not-play-dice); the tool-call risk benchmark by webofmike on dev.to. Each number was read on the original page or repository.


Written while building AgentSeism. New posts by RSS.