
We tested TypeSafe AI’s Jev to understand a simple question: where does a fast, low-cost System One model fit when the job is making a decision rather than generating text?
Jev returns structured decisions with probabilities. It’s built for choosing between defined options, not generating or reasoning through long-form responses. That makes it interesting for a different class of AI problem.
Our lab tested Jev across validation, classification, routing, extraction, ranking and prediction, then explored how those primitives could be composed into larger systems.
Here’s what we learned.
Tuning improved Kev on familiar workflows
We benchmarked Jev against open models, including Kev, across several workflows.
We then tuned Kev on 900 synthetic cases from three of those workflows.
On the workflows it had been tuned on, Kev scored 0.801 agreement with a reference AI model, compared with 0.769 for Jev.
Then we gave both models a new workflow.
Jev scored 0.655. Tuned Kev scored 0.520. The result suggests a useful trade-off.
Tuning buys depth. Jev keeps breadth.
But this was one run on synthetic data: 1,500 questions across the tuned-on workflows and 2,000 on the new workflow, with no error bars. It does not establish how either model will perform on a production workload.
That is something we would test rather than assume.
Where a model like Jev may fit
The benchmark led to a more useful question than whether Jev is “better” than another model.
What kind of decision should it be making in the first place?
A working rule emerged from our experiments: “Does the domain reward standardization? Can the output be validated easily and cheaply?”
If code can compute the answer, use code.
If the task requires open-ended reasoning, a larger model may be the better tool.
Between those two are bounded judgments: decisions where the system has a defined set of possible actions and some way to check whether the decision was right.
Classification is one example. Routing is another. So are yes/no validation, ranking and choosing the next action from a constrained set.
The point of applied research isn’t to wait for established patterns. It’s to test where the edge might be useful before the pattern exists.
That is what we explored.
Confidence can become part of the architecture
Jev returns a probability and confluence.with each decision.
In our experiments, that confidence tracked how often the model was right. That creates a useful architectural pattern.
High-confidence decisions can continue through the system.
Lower-confidence decisions can be routed to a stronger model or a person.
The important question then becomes less about whether a smaller model can replace an LLM entirely.
It becomes:
Which decisions actually need the LLM?
That distinction matters when a system is making thousands or millions of decisions.
Cost changes what is practical to evaluate
Across our lab work, we made more than 100,000 Jev requests for under $10 in total.
At that cost, a check that might otherwise be run on a sample could potentially be run on every item.
That could change how teams think about model placement.
Instead of asking one large model to handle an entire workflow, systems can potentially combine deterministic code, specialized models, frontier models and human review, using each where it makes sense.
That is a hypothesis worth testing in real operating conditions.
There are trade-offs
Jev is not a substitute for every model.
Top LLMs may achieve higher accuracy, at substantially higher cost and latency. Jev also returns a decision and probability rather than a rationale, which creates an auditability constraint. Inputs, options, probabilities, thresholds and resulting actions need to be logged if the decision has to be reconstructed later.
Hosting, language support and data requirements also affect where it can be used.
Those constraints are part of the architecture, not details to solve afterward.
What we want to test next
The experiment left us with a few questions.
How well does the breadth we observed hold up on real production data?
How stable are confidence thresholds across different kinds of decisions?
When does breaking a complex workflow into smaller judgments improve the system, and when does it allow errors to compound?
And where does a System One model create enough advantage in speed, cost or coverage to justify adding another model to the stack?
We don't have those answers yet.
But that is the useful part of testing new models early.
The goal isn't to find one model that does everything. It's to understand what each model is good enough to own.
Tell us what you're working on.


