Spend Classification with Jev: Fitting a Taxonomy into 255 Options, Rolling Up When Unsure, and What Goes to Review
Spend classification with Jev: how a category taxonomy fits its option limit, why unsure lines should roll up a level, and what stays in rules.
Written by Max Zeshut
Founder at Agentmelt
TL;DR: Spend classification — putting every invoice, purchase-order and card line into a category of your taxonomy — is the textbook job for a decision model like Jev: a closed list of options, a short piece of text per line, tens of thousands of lines a month, and a review queue whose size is the real cost. Jev picks a category with a confidence for a few thousandths of a cent per line. Two things decide whether it works. First, the taxonomy has to fit: one question accepts up to 255 options (TypeSafe's own cookbook says it is reliable to roughly 240), so a deep taxonomy is asked level by level. Second, what you do with an unsure answer: instead of sending every low-confidence line to a person, accept it one level up — "Office supplies" instead of "Printer consumables › Toner" — which keeps the spend cube useful and the review queue small. Supplier rules, amounts and dates stay in plain rules; Jev is not a calculator.
Buildable version: the spend analysis workflow — its "classify each line" step already assigns a level-1/2/3 category with a confidence score and sends the rest to review; this is the step a decision model would take over.
What spend classification with Jev is
Spend classification assigns each spend line — the description, the supplier, the GL account, the cost centre — to a category in a taxonomy, usually three levels deep, so procurement can see what the company buys, from whom, at what price. The explainer on machine learning in spend analytics covers why this used to be an annual consulting project and how classification made it monthly.
Jev, the "System One" model TypeSafe AI released in September 2026, does one part of that job: given a line and a list of categories with descriptions, it returns the category it picks, the probability it gives every other category, and a single confidence number from 0 to 1. It does not generate text, so it cannot invent a category that is not on your list — the answer is always one of your options. That makes it a natural fit for a classification step, and a poor fit for anything else in the workflow: normalising supplier names, converting currencies, comparing dates or adding up spend are code.
Step one: make the taxonomy fit
A single Jev choice question takes up to 255 options. A real spend taxonomy is bigger: UNSPSC runs to thousands of commodity codes, and even a trimmed custom taxonomy usually has a few hundred level-3 categories. There are three ways to fit it, and they combine.
| Approach | How it works | When to use it | Calls per line |
|---|---|---|---|
| Level by level | Ask level 1; then ask level 2 within the chosen level 1; then level 3 | Deep taxonomies, or when each level has fewer than ~240 options | 2–3 |
| Ask ahead, ignore the rest | In one call, ask level 1 and the level-2 question for every level-1 category; keep only the one that matches | Level 2 fits in the size budget (64k tokens per call, all questions included) | 1 |
| Classify detailed, roll up when unsure | Ask at the detailed level; if confidence is low, report the parent category instead | A flat list of up to ~240 detailed categories | 1 |
The second one sounds wasteful and is not. TypeSafe's docs call it speculative fan-out: questions are evaluated in parallel, extra questions barely change the response time, and each adds only its own tokens. The answer to "which level-2 category, assuming this is IT?" is simply ignored when the line turns out to be facilities.
Two details make any of them work better, both from TypeSafe's guidance on choice questions:
- Write descriptions that separate the options. The model sees the option names and their descriptions, not your codes. "Printer consumables: toner, ink, drums, paper for printers" beats "44103100".
- Give every level an "other" option. Without one, a line that fits nothing is forced into the closest wrong category, with a confidence that looks respectable. With one, "none of these fit" is an answer you can route.
Step two: decide what an unsure line means
This is where a decision model earns its keep. Every answer carries a confidence, and TypeSafe's own worked example shows what to do with it. In its classification cookbook, 60 company filings are classified into 75 industry groups at a 0.9 confidence cut-off: the answers above it were right 90% of the time; the answers below it were right only 40% of the time at the detailed level — but 70% of the time when reported one level up, at the parent division.
Applied to spend, that gives three outcomes per line instead of two:
| Confidence at level 3 | What happens | Why |
|---|---|---|
| High (for example ≥ 0.9) | Accept the level-3 category | Right most of the time; spot-check a sample |
| Middle | Accept the level-2 parent, flag the line | A correct "IT hardware" is more useful in the cube than a coin-flip "Laptops" |
| Low at level 2 as well, or "other" chosen | Send to a person | Genuinely unclear, or a gap in the taxonomy |
The thresholds are examples, not advice — TypeSafe's own guidance is to start conservative and tune on your own data. What the pattern changes is the size of the queue: a line the model is unsure about at level 3 is often obvious at level 2, and the category manager looking at the IT hardware total does not need to know which laptop.
Measure it yourself: after the first month, pull 200 accepted level-3 lines and 200 rolled-up lines and have someone check them. Report agreement separately for each band. If the high band is not clearly better than the middle one, the threshold is in the wrong place.
What stays in rules
TypeSafe publishes a list of Jev's known weak spots, and several of them are exactly the parts of spend data that look like judgement but are not:
- Amounts and arithmetic. "Jev is not a calculator" is TypeSafe's wording. Thresholds on value, currency conversion, totals and price variance belong in code.
- Dates and periods. Which month a line falls in, whether an invoice is inside a contract term — compare dates in code.
- Numbers as identifiers. Part numbers, GL codes and tax IDs are strings to match, not text to understand. A GL account that maps to one category should decide it outright.
- Single-category suppliers. The electricity utility and the law firm do not need a model. A supplier → category table handles a large share of lines with certainty and at no cost, and it grows every month from corrections.
- Too much context. Accuracy falls as the text sent grows with detail unrelated to the question. Send the line's description, supplier, GL account and cost centre — not the whole invoice.
The order that works: rules first (supplier map, GL map), Jev for what is left, people for what Jev is unsure about.
Learning from corrections, without fine-tuning
The spend analysis blueprint improves month over month because reviewers' corrections feed the next run. Jev cannot be fine-tuned — TypeSafe serves the same weights to every account and does not train on customer data — so corrections have to go somewhere else:
- Into the supplier table. If reviewers keep putting the same supplier in the same category, it becomes a rule, and the model never sees that supplier again.
- Into the option descriptions. When a correction reveals a boundary the model keeps missing ("software subscriptions" vs "cloud hosting"), the fix is a sentence in each description that draws the line. TypeSafe's advice on literal reading is the same: when you find yourself explaining what you really meant, that explanation is the missing half of the instruction.
- Into a pinned version. Thresholds tuned against one Jev version should stay pinned to it; when a new version ships, re-run last month's reviewed lines before moving.
Throughput and cost
A spend line is short, but the category list is not — and the list is part of what you pay for, because option descriptions are input tokens sent with every call. A level with 200 described options is a few thousand tokens per call; at $0.042 per million tokens that is about $0.0001 per line, or about $6 for 50,000 lines. Level-by-level questioning sends shorter lists and costs less; the difference is pennies either way.
Rate limits matter more than price for a month-end batch. At launch TypeSafe's limit is 1,200 requests a minute, with a warning that limits are adjusting while it adds capacity. 50,000 lines at one call each is about 42 minutes at that ceiling; three calls a line is about two hours. Schedule it after AP close and let it run.
What it costs to run as a workflow
The spend analysis workflow is a free template, or $297/month to have it run for you — one ERP source, up to 50,000 lines a month, a standard UNSPSC taxonomy, with the review sheet and the monthly opportunity briefs. It classifies with a language model today; moving that one step to a decision model is a change to the step, not a rebuild, and it should wait until a shadow test on your own lines says it is at least as good. More than one ERP, a custom taxonomy or contract-compliance checks are a custom build from $6,000. The supply chain and procurement team page lists it alongside the supplier-risk and forecasting workflows.
Questions, answered
Can Jev do transaction categorization?
Yes — picking one category from a list you define is what it is built for, and every answer comes with a confidence you can route on. The limits are practical: up to 255 options per question (about 240 reliably), text only, and no arithmetic, so large taxonomies are asked level by level and amounts stay in rules.
How many categories can Jev choose from?
Up to 255 options in one choice question; TypeSafe's own cookbook says it works reliably up to roughly 240. For a bigger taxonomy, ask level by level, or ask level 1 and every level-2 list in the same call and keep only the relevant answer.
What should happen to spend lines Jev is unsure about?
Roll them up before you send them to a person. A line that is a coin-flip at the detailed category is often clear at the parent category, and TypeSafe's worked example shows accuracy on unsure answers rising from 40% to 70% when reported one level up. Only lines that are unclear at the parent level, or where "other" was chosen, need a reviewer.
Can Jev learn from our corrections?
Not by retraining — TypeSafe does not fine-tune Jev on customer data. Corrections become supplier-to-category rules and sharper category descriptions, which is how the accuracy climbs month over month.
Sources and further reading
- Choice questions: options, descriptions and the "other" option — TypeSafe docs
- Classification using confidence (roll-up cookbook) — TypeSafe docs
- Speculative fan-out — TypeSafe docs
- Jev 1.13 known weak spots — TypeSafe docs
- Models: limits, customisation and data handling — TypeSafe docs
- Spend analysis workflow: blueprint, template and price
- Machine learning in spend analytics, explained
- Jev pricing: what one decision costs in a workflow