I wanted to try Jev. Now it’s in our product.



What I learned turning plain-language descriptions into filters on our Replenish page—and why the model’s uncertainty ended up as part of the UI instead of a reason to fail.
Jev was suddenly everywhere, and the demos made it easy to see why people were excited. I wanted to try it on a real problem inside Fabrikatör, not a toy, because that’s where you find out what happens when it almost works. My plan was modest: build it properly, test it the way I would test anything else, and see whether it earned a place in the product.
It did, and it’s now live in our app. Along the way, “passing the tests” came to mean something quite different from what I expected, and what I learned most from was what to do when the model isn’t sure.
Our Replenish page helps merchants decide what to reorder. Our customers filter it by available stock, incoming inventory, suppliers, ordering dates, Shopify tags, and other attributes. For each condition, they pick a field, a comparison (operator), and a value.
I wanted them to be able to describe what they’re looking for instead, the way people actually type:
“I want the products to have more than five incoming that I need to order in 30 days, and with the products that have Shopify tags, including clothes.”
That sentence is deliberately messy: three conditions, a relative date, a number written as a word, and no clean boundaries. The feature turns it into the filters the page already has; the user reviews the result, and the page fetches products exactly as before, with no new query path or permissions.
Jev is TypeSafe’s flagship model and the first of its System One models, built for fast, typed judgments. Instead of generating text, you give it context (state) and a set of typed questions, and it returns an answer to each, with the probabilities behind it. I used two question types: Choice (pick one of several options) and Noul (the probability that the answer to a yes/no question is yes).
That shape fit Replenish closely. Our filters already define every field and comparison, so the task was choosing among known options, not writing anything.
A structured-output LLM with enum constraints also can’t invent a field, so that isn’t the real difference. What Jev gave me was a probability for every option, which later became part of the UI; independent questions answered in parallel in one call; and low cost, from under $0.0001 to about $0.0007 per request. The models see the user’s description and our field definitions, never product records, inventory, or sales data.
My first version asked Jev to do almost everything at once. The whole filter catalogue and a set of interpretation rules went into state. One question per field picked a whole set of comparisons from up to 16 combinations. Sixteen questions counted how many values a list contained. A broad “is this request supported?” question had to score at least 0.8 before anything else counted. For my original sentence, that was 38 questions, 552 options, and 19,199 input tokens.
It failed. The support question scored 0.77, so the whole request was rejected with a generic message:

Simpler requests failed too, such as “Available stock less than 10 and Shopify status is active”. And yet all my unit tests were passing.
The tests supplied the model’s answers, so they checked what my code did with those answers, not whether Jev would give them for real language. An early browser check that looked successful had used a shortened version of my sentence. Treating that as proof was a mistake. From then on I inspected the exact requests and responses, ran the original sentences repeatedly against the live model, and built evaluation sets.
The requests showed the problem: I was asking Jev to do the catalogue’s job, holding every rule in mind, picking combinations, and counting. TypeSafe’s documentation lists counting and irrelevant state among Jev 1.13’s known weaknesses. My request had plenty of both.
So I split the work. Code does everything deterministic: it knows the catalogue, finds candidate values, builds narrow questions, and validates the result. Jev only makes small choices.
The diagram shows where the pipeline ended up; the compound steps (dashed) came a little later. The core is the same for every condition:
Here is “Available stock less than 10”, captured from the current code in our development environment, through the real application route and live model calls.¹ The code shortlisted stock_count, extracted one candidate, {v0: "10"}, and built seven questions:
field Choice over 21 fields + unsupported
compound Noul: more than one condition?
stock_count_operator Choice: eq | not_eq | lt | gt | unsupported
stock_count_eq_value Choice: v0 = "10" | none
stock_count_not_eq_value Choice: v0 = "10" | none
stock_count_lt_value Choice: v0 = "10" | none
stock_count_gt_value Choice: v0 = "10" | none
eq, not_eq and gt value questions look like lt.{
"model": "typesafe/jev-1.13",
"state": {
"condition": "Available stock less than 10",
"source": "Available stock less than 10"
},
"questions": {
"field": {
"type": "choice",
"instructions": "Which product attribute is this filter condition about?",
"criteria": {
"stock_count": "Available or on-hand inventory quantity, not incoming inventory.",
"incoming_stock": "Incoming Stock: units already on order/incoming, not available stock or units to order.",
"restock_date": "Restock Date: when to order/reorder/restock a product, not when inventory runs out.",
"unsupported": "A different field or no filter condition"
}
},
"compound": {
"type": "noul",
"instructions": "Does the request contain more than one distinct product-filter condition? Different fields or two bounds on one field are distinct conditions. A list of alternative values for one comparison is one condition. Words inside a quoted name are part of that name, not separate conditions."
},
"stock_count_operator": {
"type": "choice",
"instructions": "How should Available Stock match the requested value in `condition`? Available or on-hand inventory quantity, not incoming inventory.",
"criteria": {
"eq": "is exactly / equals one value",
"not_eq": "is not / is different from one value",
"lt": "strictly less than / before",
"gt": "strictly greater than / after",
"unsupported": "A different or missing comparison; no faithful option"
}
},
"stock_count_lt_value": {
"type": "choice",
"instructions": "Which value is specified for the 'is less than' comparison on Available Stock?",
"criteria": {
"v0": "10",
"none": "No value is specified"
}
}
}
}
The answers the code used:
| Question | Selected | Probability |
|---|---|---|
field |
stock_count (Available Stock) |
1.00 |
compound |
— | 0.05 → single condition |
stock_count_operator |
lt |
0.99 |
stock_count_lt_value |
v0 = “10” |
1.00 |
eq answer is kept.{
"answers": {
"field": {
"type": "choice",
"choice": "stock_count",
"probabilities": {
"stock_count": 1
},
"confidence": 1
},
"compound": {
"type": "noul",
"noul": 0.05
},
"stock_count_operator": {
"type": "choice",
"choice": "lt",
"probabilities": {
"lt": 0.99,
"not_eq": 0,
"unsupported": 0.01,
"eq": 0,
"gt": 0
},
"confidence": 0.99
},
"stock_count_lt_value": {
"type": "choice",
"choice": "v0",
"probabilities": {
"none": 0,
"v0": 1
},
"confidence": 0.99
},
"stock_count_eq_value": {
"type": "choice",
"choice": "v0",
"probabilities": {
"none": 0.42,
"v0": 0.58
},
"confidence": 0.17
}
},
"usage": {
"input_tokens": 1341,
"output_tokens": 416,
"cost": 5.6322e-05
}
}
The code mapped v0 back to “10”, the check approved stock_count lt 10 with every answer at 0.99 or higher, and after Apply the page used its existing filter[stock_count_lt]=10 parameter. The two Jev calls took 609 ms and 326 ms and cost $0.0000856 in total.
{
"state": {
"request": "Available stock less than 10",
"proposed_filters": [
{
"field": "Available Stock",
"comparison": "is less than",
"value": "10",
"units": "number"
}
]
},
"questions": {
"filter_0": {
"type": "choice",
"instructions": {
"question": "Does this proposed filter faithfully express a condition in the request? Check this field, comparison, complete value and units. Do not accept a weaker implied comparison.",
"proposed_filter": {
"field": "Available Stock",
"comparison": "is less than",
"value": "10",
"units": "number"
},
"field_meaning": "Available or on-hand inventory quantity, not incoming inventory."
},
"criteria": {
"faithful": "This complete filter is requested",
"different": "This filter changes or invents a requirement"
}
}
}
}
One unused answer is worth a look: stock_count_eq_value split 0.58/0.42 between “10” and “no value”, with a confidence of 0.17. That’s expected, because the description doesn’t say stock equals anything. Confidence summarises how concentrated an answer’s probabilities are. The thresholds in this post apply to the probability of the chosen option, and only on branches the code uses, so the eq branch is ignored.
With this design, my original sentence passed five command-line runs out of five and three browser runs out of three. Its interpretation request dropped from 19,199 input tokens to 2,930, and the development evaluation went from 60 to 84 exact results out of 114. The status request needed one more change: asking directly “which status value is named?” picked Active in three runs out of three, where my earlier per-value yes/no questions had missed it every time.
Compound descriptions in general were still fragile, because several conditions were interpreted from one shared text. TypeSafe’s smart-home demo has a pattern for this: a Noul detects compound requests, an LLM splits them, and TypeSafe evaluates each part.
I followed it. The first call now also asks whether the description contains more than one condition. Below 0.5, the code uses that call’s answers directly. Above 0.5, a small general-purpose LLM (gpt-4.1-mini via OpenRouter) splits the description into single conditions, each with the exact excerpt it came from. It’s the only generative step, and it never chooses fields or comparisons. If any excerpt isn’t found verbatim, the code discards the split, reads the whole sentence as one condition, and tells the user the result may be incomplete. Otherwise each part goes through the same small questions, up to four at a time.
Cutting at “and” wouldn’t work: Supplier contains "Acme and Sons" has to stay one condition. Splitting wasn’t the main fix, though; in my probes, splitting alone left the old status question wrong in three runs out of three. It isn’t a guarantee either: Jev reads the splitter’s restatement of each part, and a split can drop a condition. Values still come from the user’s words and the check compares every filter with the original sentence, but that can’t catch everything, which is one reason I count confident wrong filters separately.
Then another request made me question the product behaviour: “Available stock less than 10 and Shopify status is available.”
Jev understood the stock condition. The problem was “available”, because our Shopify Status filter only has Active, Draft, and Unlisted. That part didn’t clear its threshold, so the app returned no filters at all and asked the user to be more specific. That felt wrong. I had the stock condition, the other part was clearly about Shopify Status, and Jev had already said how likely each status value was. Why make the user start again?
My implementation treated confidence as a gate: every answer had to clear its threshold, or the user got a failure message. So I changed the response. The stock filter stays in the proposal, and the status condition becomes a question that quotes the unclear words and offers Active, Draft, and Unlisted, most likely first. The user can pick one, enter a value with the existing filter editor, or remove the condition.

The same change helped when a filter looked right but the check wasn’t confident enough: the app now asks for confirmation instead of discarding it. Every condition the app identifies ends up as a ready filter or an unresolved item, and if a condition might have been missed, the review says so. Only two kinds of request are still refused outright: “this OR that” across different fields, or other mixed logic our filters can’t express, and requests to change data rather than view it.
The ranked options come from probabilities Jev had already returned, so they cost no extra call, and the change didn’t lower any thresholds (0.7 for field, comparison, and value choices, with narrower rules for list members and single candidates; 0.8 for request-level checks). What changed is what happens below them.
Partial results also settled a UI question. For a while I had tried applying generated filters immediately. It saved a click, but the filters, tab, product selection, and table all changed at almost the same moment, and applying only the understood conditions could quietly change what a request meant. I went back to a compact review step: understood and unresolved conditions sit side by side, and Apply stays disabled until every open question is answered or removed. Example prompts stay visible, and corrections reuse the filter controls already on the page.
Some ambiguity isn’t a model problem at all. “Tags include clothes and discount” could mean either tag or both, and no probability tells you which your product should do. Our filter supports “contains any of”, so I decided a list means any of; explicit “both” or “all” wording shouldn’t be applied confidently, and the check is told to flag it. The model can tell you what the words probably mean; you still have to decide what your product supports.

By the time it went live, “passing the tests” meant three things:
In the evaluation run I recorded, the development set scored 104 of 116 exact matches and a fresh held-out set 15 of 20. The number I watch most closely is confident wrong filters: cases where the app proposes a filter that changes what the user asked for without flagging any doubt. There were none on the development set and one on the held-out set. A miss that turns into a question is an inconvenience. A confident wrong filter shows the user the wrong products, and they might not notice.
It isn’t finished. I set the thresholds by hand on the development set; they aren’t calibrated, and results near them vary between runs (one compound phrasing scored 0.48 against the 0.5 cutoff). The goal is still at least 95% on a representative set of real-world descriptions. If Jev or the splitter fails or times out, the panel says describing isn’t available right now, and the user’s current filters stay as they were.
So I also record what people do in review: apply or cancel, which suggested option they pick, whether they enter their own value or remove a condition, and whether they edit the proposal. An edit can reflect a changed mind, and applying unchanged doesn’t prove the proposal was right, but these signals give me real descriptions to learn from. Turning them into evaluation cases is the next step.
I started out wanting to try Jev on a real problem. Building it into our own product felt very different from watching the demos online: I had to decide where it belongs, what the user sees, and what happens when it’s almost right. Now that it works on Replenish, I’m excited about the rest of the app. The same engine can serve any page that uses our filters, and there are plenty of places where describing what you want beats clicking through menus.
How this post is made: I built the feature with AI coding agents working under my supervision. I set the direction, reviewed the work, tested it, and made the product decisions; the agents did the implementation. Where this post says “I built” or “I changed”, that is how it worked. Everything in the post comes from that work: the failures, numbers, screenshots, and request data are from real development sessions and live model calls, and none of it is mocked. I also used an AI assistant (Claude & Codex) as my editor for the post itself: it helped restructure the draft, check claims against our code and TypeSafe’s documentation, draw the diagram, and record and trim the videos. The decisions, and any mistakes, are mine.
¹ None of the model responses are mocked. The JSON examples are trimmed excerpts of the captured requests and responses; authentication and transport headers are omitted. The screenshots are from earlier development builds and are not the network trace for this request.