Lead Scoring with Jev: One Question per Criterion, Weights in Your Rules, and a Confidence Field in the CRM
How to use Jev for lead scoring: one question per criterion, weights kept in your rules, confidence written to the CRM, and what a model should not touch.
Written by Max Zeshut
Founder at Agentmelt
TL;DR: Lead scoring with Jev means splitting "is this lead any good?" into separate, narrow questions — how specific the need is, what timeline they state, who they are in the decision, what budget they mention — and letting Jev answer each one from the enquiry text with a score and a confidence, in well under a second, for a fraction of a cent. The weights that turn four answers into one lead score stay in your rules, where sales ops can read and change them; firmographic fit stays in rules too, because it is lookup, not judgement. The confidence goes into the CRM beside the score, so "hot" means "hot and the model was sure", and anything it was unsure about is treated as not yet known — which is exactly when the first reply should ask the qualifying question. Jev decides where the lead goes; it does not write to the lead, look anything up or do arithmetic.
Buildable version: the inbound lead qualification workflow — it already splits scoring into deterministic fit rules plus a model reading the message, which is the slot a decision model fills.
What lead scoring with Jev is
Jev is a "System One" model from TypeSafe AI: you give it a piece of text and a set of typed questions, and it returns typed answers with probabilities — a choice from options you listed, a score on levels you described, or a yes/no probability — instead of generated text. Used for lead scoring, the text is the enquiry (the form message, the email, the chat transcript) and the questions are your qualifying criteria. It answers all of them in one call, in parallel, typically in 70–500 milliseconds, and at $0.042 per million input tokens a lead costs a few thousandths of a cent to score. The Jev pricing breakdown has the arithmetic.
What it does not do matters as much. Jev does not enrich — it cannot look up company size or funding. It does not count page visits or work out that "next Tuesday" is within 14 days; TypeSafe's own list of known weak spots says to keep arithmetic and date comparison in code. And it does not draft the reply; that stays with a language model or a template. In a lead workflow it is the judgement in the middle: read what the person wrote, answer narrow questions about it, hand the answers to rules.
Why one question per criterion
The tempting shortcut is one question: "How qualified is this lead, 1 to 5?" It fails for the same reason a single-number health score fails — "3" hides whether the problem is budget, timing or fit — and TypeSafe's guidance on scores says as much: keep each question to one dimension, because an input that is high on one thing and low on another cannot be placed, and the confidence drops. Split it, and combine in code.
Two more rules from the same guidance shape how the questions are written:
- Describe situations, not degrees. "Names a specific problem and what it is costing them" gives the model something to match; "moderately interested" does not. Each level is judged on its own — the model does not see the level's number or its neighbours — so every level has to stand alone as a description.
- Use the right question type. Ordered levels (how specific is the need) are a Score. Categories with no order (who is this person in the decision) are a Choice, with a "not stated" option so a missing answer is reported rather than guessed. A plain gate (is this a sales enquiry at all?) is a yes/no question.
The questions, criterion by criterion
The classic four are budget, authority, need and timeline. Here is how each becomes a question Jev can answer well, and which ones should not be asked of a model at all.
| Criterion | Question type | Options or levels (illustrative) | Why this type |
|---|---|---|---|
| Is it a sales enquiry? | Yes/no | — | Filters support requests, job applications and spam before anything is scored |
| Need | Score, 4 levels | No problem described · general interest · a specific problem named · a specific problem with a cost or deadline attached | Ordered: more specific is better, and the top level is the one that predicts a meeting |
| Timeline | Choice | This month · this quarter · later · no timeline stated | Asks what they said, not what date it is — date arithmetic stays in code |
| Authority | Choice | Decides · evaluates for a team · researching for someone else · not stated | Categories, not a scale; "not stated" is common and should not be scored as low |
| Budget | Choice, or a rule | Bands as stated · not stated — or read the form's budget field directly | If the form asks for budget, use the field; a model adds nothing to a number already given |
| Company fit | Rule, not a model | Size band, industry, region from enrichment | Lookup, not judgement; rules are transparent and sales ops owns them |
Asked together, those five model questions are one call. The response carries, for each Score and Choice, the answer, the probabilities across the options, and a confidence from 0 to 1; the yes/no question carries a probability (and no separate confidence field — the probability is the signal).
The weights stay in your rules
Jev's answers are inputs to your lead score, not the lead score. Combining them is a few lines of rules that sales ops can read — the decision engine pattern: the model reads and proposes, the rules decide.
An illustrative version:
- Normalise each score to 0–1. A four-level scale returns 0 to 3, so divide by 3; TypeSafe's docs make the same point, so that a top score on a short scale does not outweigh a top score on a long one.
- Map each choice to points. "Decides" 1.0, "evaluates for a team" 0.7, "researching for someone else" 0.3, "not stated" 0.5 — neutral, not zero, because silence is not a no.
- Weight and add. Need 0.4, timeline 0.3, authority 0.2, budget 0.1 — or whatever your won deals say matters. The weights are yours; when the result disagrees with what your best rep would have decided, change them and re-run last month's leads.
- Multiply by fit. An excellent enquiry from outside your market is still outside your market; the rule-based fit score gates the intent score.
- Band it. Hot, warm, cold — thresholds in a sheet, not in code, so they can be moved without a deploy.
Because every piece is visible, "why was this lead hot?" has an answer a rep will accept: need 3 of 3, timeline this month, fit A.
Where the confidence goes
This is the part that makes a decision model different from asking an LLM for a score. Each answer comes with a confidence, and TypeSafe's recommended pattern is three ranges: act when it is high, proceed with caution in the middle, do not act when it is low — with thresholds that scale with what a wrong answer costs.
For leads, the costs are lopsided, and the thresholds should be too:
| Wrong answer | What it costs | Threshold it deserves |
|---|---|---|
| A lead marked hot that isn't | Ten minutes of a rep's time | Moderate: hot requires the need and timeline answers above, say, 0.8 confidence |
| A lead marked cold that isn't | Possibly the deal — it went to a competitor who replied | Strict the other way: low confidence never makes a lead cold, it makes it warm |
| A criterion the model was unsure about | Nothing, if you treat it as unknown | Below 0.5, record "not stated" and let the reply ask |
Then write it all back: the lead score, the band, each criterion's answer, and the lowest confidence of the four, into CRM fields. Two things follow. Reps can filter to "hot, confident" and trust it. And after a month you can compare bands against meetings booked — which is the only number that tells you whether the thresholds are right.
Measure it yourself: before switching anything on, take last quarter's inbound leads, run the questions over them, and compare the bands with which leads became meetings. The shadow test for Jev's accuracy walks through doing that without trusting anyone's benchmark, including TypeSafe's.
What to watch for in sales text specifically
- Enquiries that argue for themselves. "URGENT — enterprise deal, decision this week" is text written to be read a certain way. TypeSafe lists adversarial or steering content as a known weak spot; be explicit in the level descriptions (for example, "a stated deadline tied to a named event") and test a few pushy enquiries before trusting the top band.
- Literal reading. Jev answers the question you wrote, not the one you meant. If "need" should exclude people describing a competitor's problem, say so in the level description.
- Language. English is where accuracy is best; if a share of your leads write in other languages, test those separately.
- Speed is not the bottleneck anymore. A decision in half a second means scoring never delays the first reply. The slow step is drafting the reply with an LLM — which is the right place to spend the time, because that is the part the lead reads.
What it costs
The model line is close to nothing: an enquiry of around 1,000 tokens with five questions is about $0.00004, so 1,000 leads a month is about four cents. What you are paying for is the workflow around it. The inbound lead qualification workflow is a free template, or $247/month to have it run (up to 1,000 leads a month, one CRM, one enrichment provider); multi-territory or multi-product routing is a custom build from $3,000. Solo agents and brokers usually want the lighter speed-to-lead workflow instead, from $197/month. The sales team page lists both with the outbound and expansion workflows they pair with.
Questions, answered
Can Jev do lead scoring?
Yes, for the part of lead scoring that is judgement over text: how specific the need is, what timeline and role the person states, whether it is a sales enquiry at all. It returns a score or a choice with a confidence for each. Firmographic fit, page-visit counts and date arithmetic belong in rules, and the reply belongs to a language model or a template.
How do I weight lead scoring criteria with Jev?
Outside the model. Ask one question per criterion, normalise each answer to 0–1, and weight them in your own rules — a sheet sales ops can edit. Jev's documentation recommends the same: split a complex judgement into single-dimension questions and combine them in code.
Is Jev better than an LLM for lead scoring?
Faster and far cheaper, and each answer comes with a confidence an LLM does not give reliably. Whether it is as accurate on your leads is something to test on last quarter's data before switching; on TypeSafe's own evals it sits with the mid-tier LLMs, a few points below the best.
Does Jev replace the lead score in my CRM?
No — it feeds it. The CRM keeps the score, the band and the history; Jev supplies the answers the score is computed from, plus the confidence you store beside it.
Sources and further reading
- Score questions: writing levels and combining scores — TypeSafe docs
- Confidence: three ranges and thresholds that scale with risk — TypeSafe docs
- Jev 1.13 known weak spots — TypeSafe docs
- Introducing System One Models & Jev — TypeSafe AI
- Inbound lead qualification workflow: blueprint, template and price
- Speed-to-lead workflow
- Decision engines for agentic AI, explained
- How AI agents qualify real estate leads