/

News

TypeSafe’s JEV

TypeSafe's Jev answers preset questions in milliseconds with confidence scores. Cheap, fast, bounded decisions could put AI inside the runtime continuously.

TypeSafe's Jev answers preset questions in milliseconds with confidence scores. Cheap, fast, bounded decisions could put AI inside the runtime continuously.

If AI can make thousands of decisions a second, the architecture starts to change.

Ex-OpenAI researcher Diogo Almeida’s TypeSafe just emerged from stealth with Jev, a model that answers preset questions inside software with confidence scores attached. 

TypeSafe says Jev runs at $42 per billion input tokens, with output free, and returns responses in roughly 70 to 500 milliseconds. Because it only chooses from defined outputs, it also avoids the kind of open-ended hallucination you get from generative models.

Our engineering channel lit up almost immediately. The Doom demo probably helped.

Jev was making repeated decisions fast enough to control the game in real time. That is a useful demonstration of the workload it is built for, tight feedback loops where software needs a fast answer, over and over again.

A lot of production AI uses LLMs for work that is much more constrained than generation. Route this request. Classify this document. Score this record. Decide whether this needs review. Pick which model should handle the next step.

You can use a frontier LLM for all of that. But once those decisions are happening thousands or millions of times concurrently, or over a short period, the economics start to shape the product. 

A few seconds of latency and a small inference cost look very different at high volume. So does asking a generative model for an answer and then turning that answer back into something software can reliably use.

If Jev’s numbers hold up in production, it could make a different class of real-time AI practical.

Continuous model routing. Real-time document processing. Security analysis on live traffic. Decisioning inside claims or support workflows. Fast checks around agents before they take an action.

That is where the economics get interesting for enterprise teams.

The bigger shift is being able to make thousands of structured, discrete decisions a second. At that point, AI can sit inside the runtime continuously, routing, checking, scoring and responding as the system changes. That only becomes practical when the cost and latency are low enough.

That also pushes teams toward a more deliberate architecture.

Use language models where you need reasoning or generation. Use System One or System Two models where the job is a fast, bounded decision. Use deterministic software where the rules are already known.

The engineering challenge is choosing the right tool for each part of the system. This is another example of what our engineers have said all along, the engineering challenge was, is, and will be choosing the right tool for each part of the system.

There is still a lot to prove before enterprise teams put something like Jev into critical workflows, especially around reliability, security, data handling, deployment, monitoring, confidence thresholds and failure behaviour. 

We’re looking forward to seeing where this goes as the technology matures and more of these use cases become practical in production.

And yes, Doom was cool.

Have an AI priority that has to work?

Have an AI priority that has to work?

Tell us what you're working on.

Senior AI expertise for work that has to perform.

Senior product and engineering judgment

for AI that has to work.

Senior AI expertise for work that has to perform.