Is Jev Accurate? What "Can't Hallucinate" Really Means, What the Tests Show, and a Shadow Test to Run Before You Trust It
Is Jev accurate? What "can't hallucinate" means, what vendor and independent tests show, and a shadow test an operations team can run before trusting it.
Written by Max Zeshut
Founder at Agentmelt
TL;DR: Jev is accurate enough for a lot of routing and classification work, a few points less accurate than the best LLMs, and its confidence scores are useful but not honest enough out of the box to trust without checking. "Can't hallucinate" means something narrower than it sounds: every answer fits the format you defined — TypeSafe itself says its 0% hallucination figure "is not empirical" — but a well-formed answer can still be the wrong one. The evidence so far points one way: accuracy depends less on the model than on how the question is asked and what data it is asked about. One independent tester took a phishing check from 62.6% to 95.0% by splitting one question into five. So the only accuracy number that matters is the one you measure on your own past cases, in a shadow test, before anything acts on the answers. It takes an afternoon of someone's time and about five cents of model cost.
Buildable version: the classification and triage workflows on this site gate each model decision on a confidence score or a review step — the shadow test below is how to set that score for your data, whichever model answers.
Is Jev accurate?
Jev is a "System One" model from TypeSafe AI: it reads text and answers typed questions — pick one of these options, place this on that scale, is this statement true — with probabilities and a confidence. Asked "is it accurate", the honest answer has three parts.
- Compared with LLMs, it is in the middle of the pack. On TypeSafe's own four workflow evals Jev scores 67.8% agreement with a frontier-model reference, level with Sonnet 5 and GPT-5.6 Terra and five to six points behind the best two. In OpenRouter's independent test on 3,080 banking support messages, it scored 81.0% against Claude Opus 5's 84.4%.
- Its confidence sorts right from wrong well — on average. In the same OpenRouter test, answers at 0.99 confidence or above were right 96.3% of the time, and answers below 0.5 only 29.6%.
- But averages hide the questions where it is confidently wrong. In a pre-registered test by Rajesh Beri, one internal policy question came back right 44.7% of the time while Jev put an average probability of 0.74 on its answers.
None of that is unusual for a model. What is unusual is the marketing line, so start there.
What "can't hallucinate" means, and what it doesn't
An LLM produces text, so it can produce anything: a category that is not on your list, a field that does not exist, a confident paragraph about a policy you never wrote. Jev produces none of that. Its answers are drawn from the options you defined, so an answer that does not fit your schema is, as TypeSafe puts it, mathematically impossible. That is real, and for automation it matters — a step deep inside a workflow cannot choke on a malformed answer.
But look at how TypeSafe itself footnotes the 0% hallucination figure in its launch post: the number is "not empirical" — it is there because schema matching is guaranteed. The other half of what people mean by hallucination, a confident answer that is wrong, is simply called an error in Jev's case, and Jev makes those at roughly the rate the accuracy numbers above suggest. Its documentation is explicit that calibration is measured across groups of predictions and "does not guarantee" that any individual answer is correct.
The practical translation: you no longer need to validate the shape of what comes back. You still need to validate whether it is right, and the confidence tells you where to look.
What the tests show so far
| Test | Who ran it | Task | Jev | Comparison | What it tells you |
|---|---|---|---|---|---|
| Workflow evals | TypeSafe (vendor) | 4 business workflows, frontier-model answers as reference | 67.8% | Opus 5 73.1%, Sol 74.1%, Sonnet 5 67.8% | Mid-tier LLM agreement at a fraction of the cost; weakest on invoices (61.8% vs 79.1%) |
| Banking77 intents | OpenRouter (independent) | 3,080 messages, 77 intents | 81.0% | Opus 5 84.4% | A 3.3-point gap; high-confidence answers 96.3% right |
| Calibration study | Rajesh Beri (independent, pre-registered) | 5,721 calls across 21 experiments | Calibration error 0.107 on unfamiliar data, 4.4x the noise floor | 0.024–0.032 on a public benchmark | Confidence is better on data like the public benchmarks than on yours |
| Phishing, one question | Same study | 1,000 test emails | 62.6% | Claude Haiku 81.3% | One broad question is a bad question |
| Phishing, five questions | Same study | Five narrow questions + a simple model on 1,000 labels | 95.0% | Haiku, same setup, 93.2% | Decomposition beats model choice |
Two patterns run through all of it. First, decomposition matters more than the model: TypeSafe's evals found every LLM more accurate when the task was split into narrow questions than when it was given as one prompt, and the phishing result shows the same for Jev. Second, confidence needs calibrating per question: Beri found Jev's yes/no answers under-confident and its choice and score answers over-confident, by different amounts per question. A threshold that is right for one question is not right for the next — which TypeSafe's own documentation also warns: don't carry a threshold tuned on one question type over to another.
Where Jev is weak, by its maker's own account
TypeSafe publishes a list of known weak spots for the current version. For an operations team, the ones that matter:
- Literal reading. It answers the question as written, not as meant. Negations and implied conditions are read at face value.
- Numbers and dates. It does not count reliably and reads dates as text, so "within 30 days" and "over $5,000" belong in rules.
- Too much context. Accuracy falls when the text sent is full of material unrelated to the question.
- Adversarial text. Content written to steer the answer — a phishing email, an enquiry arguing it is urgent — can move it.
- Language. English is where it is best, per TypeSafe's models page; other languages are handled, not equally well.
Every one of those is fixable in the workflow rather than the model: rules for the arithmetic, filtering before the question, narrower wording, a person on the adversarial steps.
The shadow test: how to find your own accuracy
A shadow test runs the model on real past work without letting it act, then compares its answers with what your team actually decided. It is the same discipline as testing any agent before launch, made cheap by a model that costs a hundredth of a cent a call.
| Step | What you do | What you need |
|---|---|---|
| 1. Pick one decision | Routing tickets, scoring leads, classifying spend lines — one at a time | The questions, written narrowly: one dimension each, options with descriptions, an "other" |
| 2. Gather the answer key | 300–500 recent cases where you know the right answer: the queue a ticket ended in, the category a reviewer confirmed | An export; include the awkward cases, not just the clean ones |
| 3. Pin the version | Use the exact version name, not "latest", which moves when a new release ships | One setting |
| 4. Run it in shadow | Every case through Jev; store the answer, the probabilities and the confidence | About $0.05 of model cost for 500 cases |
| 5. Find the band | On half the cases, find the confidence above which accuracy clears your bar, per question; check it holds on the other half | A spreadsheet |
| 6. Read the misses | Sort every wrong answer: badly worded question, missing context, genuinely ambiguous, or a known weak spot | An afternoon |
| 7. Price the threshold | Share above it × error rate there × cost of a wrong answer, against share below × minutes of review | The cost of one mistake, honestly estimated |
| 8. Go live, keep the log | Act above the threshold, review below it, and re-run the same set whenever the version changes | The log you already keep |
Step 6 is where most of the gain is. When a wrong answer makes you think "that's not what I meant", the fix is in the question, not the model — TypeSafe's advice is that your explanation of what you meant is the missing half of the instruction. Rewrite, re-run, and the band usually moves.
Measure it yourself — the benchmark that matters: before judging the model, have two of your people answer the same 100 cases independently and count how often they agree. If they agree 90% of the time, a model that is 96% right above its threshold is already more consistent than a second reviewer on those cases — and the cases below the threshold are the ones your people disagreed on too.
Beyond accuracy: what else to check before relying on it
- Data terms. Zero data retention is offered to enterprise customers; TypeSafe says its service is currently based on the US West Coast. Check both against what your data allows.
- Moving versions. The "latest" name changes with each release, and answers can shift with it. Pin, and re-test before moving.
- Rate limits. At launch, 1,200 requests a minute, with a warning that limits are adjusting while capacity is added. Fine for most workflows; plan month-end batches around it.
- One vendor, closed weights. Keep a tested fallback to an LLM for the same questions. Community "Jev-style" open models exist, but check the licence: OpenJev, for one, says it is not affiliated with TypeSafe and its weights are licensed for non-commercial use only.
None of these is a reason not to use it. All of them are reasons to design the step so a model can be swapped and a person can take over — which is how the human-in-the-loop steps in every workflow here are already built.
What it costs
The shadow test itself costs almost nothing in model time — 500 cases is a few cents — and an afternoon of a person's time, which is the real cost and the one worth paying. The classification and triage workflows here are a free template each, or run for you from $197–$297/month; where a model decides, the confidence threshold is the number this test sets. The comparison of which steps suit Jev and which suit an LLM covers where it is worth testing at all, and the pricing breakdown shows why the review queue, not the model, is the line to optimise.
Questions, answered
Is Jev accurate?
Accurate enough for many routing and classification jobs, and a few points below the best LLMs: 67.8% against 73–74% on TypeSafe's own workflow evals, and 81.0% against Claude Opus 5's 84.4% on an independent banking-intent test. Its high-confidence answers are much more accurate than its low-confidence ones, which is what makes it usable — but measure it on your own cases before trusting a threshold.
Can Jev hallucinate?
It cannot return an answer outside the options you defined, so it never produces an invented category or a malformed result. It can still pick the wrong option. TypeSafe notes that its 0% hallucination figure is not an empirical measurement but a consequence of the format guarantee.
Are Jev's confidence scores reliable?
On average they sort right from wrong well, but they need checking per question. An independent study found yes/no answers under-confident and choice and score answers over-confident, and one question where Jev was right 44.7% of the time at an average probability of 0.74. Set thresholds from your own shadow test, not from a default.
How do I test Jev before using it?
Run it in shadow on 300–500 past cases where you know the right answer, with the model version pinned. Find the confidence band where accuracy clears your bar on half the cases, check it on the other half, read every miss, and re-run the set whenever the version changes.
Sources and further reading
- Introducing System One Models & Jev — TypeSafe AI (see the "Hallucination and type-safety" notes)
- System One: calibration and its limits — TypeSafe docs
- Confidence — TypeSafe docs
- Jev 1.13 known weak spots — TypeSafe docs
- Workflow evals — TypeSafe
- Jev vs Claude Opus 5 on classification — OpenRouter (22 Sep 2026)
- TypeSafe Jev: calibration, decomposition and a shadow eval — Rajesh Beri (20 Sep 2026)
- OpenJev model card — Hugging Face
- How to evaluate and test AI agents before production
- Human-in-the-loop for AI agents