/

Labs

Jev’s edge is decision breadth

TextLayer tested TypeSafe AI’s Jev across classification, routing, validation, ranking and prediction to understand where fast, low-cost decision models fit alongside code, LLMs and human judgment.

TextLayer tested TypeSafe AI’s Jev across classification, routing, validation, ranking and prediction to understand where fast, low-cost decision models fit alongside code, LLMs and human judgment.

Gareth Sharpe

AI Architect

We tested TypeSafe AI’s Jev to understand a simple question: where does a fast, low-cost System One model fit when the job is making a decision rather than generating text?

Jev returns structured decisions with probabilities. It’s built for choosing between defined options, not generating or reasoning through long-form responses. That makes it interesting for a different class of AI problem.

Our lab tested Jev across validation, classification, routing, extraction, ranking and prediction, then explored how those primitives could be composed into larger systems.

Here’s what we learned. 

Tuning improved Kev on familiar workflows

We benchmarked Jev against open models, including Kev, across several workflows.

We then tuned Kev on 900 synthetic cases from three of those workflows.

On the workflows it had been tuned on, Kev scored 0.801 agreement with a reference AI model, compared with 0.769 for Jev.

Then we gave both models a new workflow.

Jev scored 0.655. Tuned Kev scored 0.520. The result suggests a useful trade-off.


Benchmark result: Kev matched Jev only where tuned
Workflows Kev was tuned on
900 synthetic cases from three workflows
Jev
0.769
Kev, untuned
0.626
Kev, tuned
0.801
0
0.5
1.0
New workflow
2,000 synthetic cases
Jev
0.655
Kev, untuned
0.467
Kev, tuned
0.520
0
0.5
1.0


Tuning buys depth. Jev keeps breadth.

But this was one run on synthetic data: 1,500 questions across the tuned-on workflows and 2,000 on the new workflow, with no error bars. It does not establish how either model will perform on a production workload.

That is something we would test rather than assume.

Where a model like Jev may fit

The benchmark led to a more useful question than whether Jev is “better” than another model.

What kind of decision should it be making in the first place?

A working rule emerged from our experiments: “Does the domain reward standardization? Can the output be validated easily and cheaply?”

If code can compute the answer, use code.

If the task requires open-ended reasoning, a larger model may be the better tool.

Between those two are bounded judgments: decisions where the system has a defined set of possible actions and some way to check whether the decision was right.


The right tool for the decision: Different decisions call for different tools
Use the cheapest reliable decision-maker for the job. If code can compute the answer, use code. Escalate as decisions become less bounded or higher risk.
1. Computable
Code
If it can be calculated, calculate it.
Examples
Math and logic
Deterministic rules
Data transforms
2. Bounded judgment
Jev / specialized model
Choose, classify, rank, route or validate.
Examples
Classification
Routing and triage
Yes/no validation
Ranking and next best action
3. Open-ended reasoning
LLM
Interpret, synthesize, and reason.
Examples
Open-ended analysis
Complex judgment
Multi-step reasoning
Content generation
4. High-risk / uncertain
Human
Review when judgment matters.
Examples
Ambiguous or sensitive decisions
Edge cases
Regulated workflows
High impact outcomes


Classification is one example. Routing is another. So are yes/no validation, ranking and choosing the next action from a constrained set.

The point of applied research isn’t to wait for established patterns. It’s to test where the edge might be useful before the pattern exists.

That is what we explored.

Confidence can become part of the architecture

Jev returns a probability and confluence.with each decision.

In our experiments, that confidence tracked how often the model was right. That creates a useful architectural pattern.

High-confidence decisions can continue through the system.

Lower-confidence decisions can be routed to a stronger model or a person.

The important question then becomes less about whether a smaller model can replace an LLM entirely.

It becomes:

Which decisions actually need the LLM?

That distinction matters when a system is making thousands or millions of decisions.


Confidence routing: Use Jev for fast, bounded decisions. Escalate the rest.
A simple pattern for combining speed, cost and accuracy.
Input
Options, context, structured data
Jev · System one
Structured decision + confidence score
Confidence?
High confidence
Take action
Automate, apply, or store result
Examples
Classify
Route
Validate
Extract
Rank
Low confidence
Stronger model or human review
Handle uncertain or complex cases
Examples
Send to LLM
Request more context
Human review
Multi-step reasoning

Cost changes what is practical to evaluate

Across our lab work, we made more than 100,000 Jev requests for under $10 in total.

At that cost, a check that might otherwise be run on a sample could potentially be run on every item.

That could change how teams think about model placement.

Instead of asking one large model to handle an entire workflow, systems can potentially combine deterministic code, specialized models, frontier models and human review, using each where it makes sense.

That is a hypothesis worth testing in real operating conditions.

There are trade-offs

Jev is not a substitute for every model.

Top LLMs may achieve higher accuracy, at substantially higher cost and latency. Jev also returns a decision and probability rather than a rationale, which creates an auditability constraint. Inputs, options, probabilities, thresholds and resulting actions need to be logged if the decision has to be reconstructed later.

Hosting, language support and data requirements also affect where it can be used.

Those constraints are part of the architecture, not details to solve afterward.

What we want to test next

The experiment left us with a few questions.

How well does the breadth we observed hold up on real production data?

How stable are confidence thresholds across different kinds of decisions?

When does breaking a complex workflow into smaller judgments improve the system, and when does it allow errors to compound?

And where does a System One model create enough advantage in speed, cost or coverage to justify adding another model to the stack?

We don't have those answers yet.

But that is the useful part of testing new models early.

The goal isn't to find one model that does everything. It's to understand what each model is good enough to own.


Have an AI priority that has to work?

Have an AI priority that has to work?

Tell us what you're working on.

Senior AI expertise for work that has to perform.

Senior AI expertise for work that has to perform.

Senior product and engineering judgment

for AI that has to work.