Jev explained

7 min read

When I first heared about "Jev" and my first reaction was: "is this another LLM?"

I used it in Agent Battle Royale to pick which specialist should answer a prompt. I also used it as a judgment layer when a deploy goes wrong.

and the short answer. No.

What is Jev?

Jev is TypeSafe AI's first System One model.

System One is the fast judgment layer which means the snap/fast decision.

System Two is the slow, verbal thinking. LLMs behave more like System Two. LLMs think, write multi-steps plan, chat, execute etc.

Jev does the opposite job (system one)

Jev does not generate text, images, or chat. It returns typed values, probabilities, and a confidence score your code can use to pick the next step.

Your application reads it like a function call.

The name Jev comes from the Jevons, the economist. He has quoted " Cheaper intelligence, more use" . TypeSafe also prices output as free, which is why people keep saying Jev is about saving money, and leading to more usage.

State and typed questions

In an LLM you send a prompt a short or long one. Then you parse the reply, check for the errors, and if all good then you can use it in your application.

In Jev you send two things:

  1. State. This is the facts. Could be a Ticket text, incident notes, a user prompt, a JSON object. This is the Context the model will judge.

  2. Typed questions. Also called primitives. You decide the shape of the answer before you call the model.

Hand-drawn diagram of state and typed questions going into Jev

There are three primitives:

  1. Choice. Pick one option from a list you defined. You get the winner, a probability for each option, and a confidence.

  2. Score. Rate on an ordered scale. You get a number, a distribution, and a confidence.

  3. Noul. Yes or no, returned as a probability from 0 to 1.

Hand-drawn cards for Choice, Score, and Noul

You lock the answers. Jev cannot return an option you did not list.That is why people say it cannot hallucinate a type. It can still be wrong. It just cannot invent a fourth team called "magic".

I used this on a deploy incident in a LinkedIn post. Checkout page starts failing minutes after a release. Instead of asking an LLM to write a novel, I asked Jev three questions:

  1. Choice: what is the likely cause?

  2. Score: how severe is it?

  3. Noul: should we roll back?

72% on the deployment as the cause. Score 1.45 out of 2. 81% yes on rollback.

Probability and confidence are not the same number.

This is one more thing people mix up. The 72% belongs to one option. Confidence says how clearly that option stands out from the rest. It does not mean Jev is right 72% of the time. The engineer still checks the evidence.

LLMs vs Jev

Maybe you are thinking: I already get JSON from an LLM. Why do I need this?

LLM Jev
Job Write, chat, code, explain Judge, route, score, gate
You send A prompt State and typed questions
You get Text you must parse Values your code can use
Speed Seconds is normal Tens to hundreds of ms
Output cost You pay for tokens out Output is free
Can it invent an option? Yes No, not outside your list

Use an LLM when a human needs words. Use Jev when software needs a decision.

I had a custom router in my Agentic AI application. Prompt in, I decide who should speak, then an LLM writes the answer. That decision layer was the expensive, fiddly bit. I replaced it with Jev.

n8n still manages the events. The specialist still talks. Jev only picks who talks.

Cost and latency

Who does not like saving money? Tokkemaxxing is no longer a thing to flex.

TypeSafe bills input tokens. Output is free. They call it too cheap to meter. Where LLMs return a paragraph, Jev just returns free value, probilities etc. What bill is just "input" and that too very cheap.

They also claim the call lands in about 70 to 500 ms. I will not pretend I independently benchmarked every model on earth. What I felt in my own demos: it is fast enough to sit in a UI and show the routing live, which an LLM judge was not.

If your agent spends money deciding who should speak, that layer is the wrong place for a chat model.

Code demo

Here is the shape. One state, three questions. Same idea as the incident example.

{
  "model": "jev-latest",
  "state": {
    "incident": "Checkout started failing 4 minutes after deploy. Error rate jumped. Payments service is healthy.",
    "last_change": "web checkout release"
  },
  "questions": {
    "cause": {
      "type": "choice",
      "instructions": "What is the likely cause?",
      "criteria": {
        "deploy": "The new release broke checkout",
        "payments": "The payments provider is down",
        "traffic": "A traffic spike, not a code issue"
      }
    },
    "severity": {
      "type": "score",
      "instructions": "How severe is this for customers?",
      "criteria": ["low", "medium", "high"]
    },
    "rollback": {
      "type": "noul",
      "instructions": "Should we roll back the release now?"
    }
  }
}

What comes back is keyed the same way. cause has a winner and probabilities. severity has a score. rollback has a yes/no probability.

In Agent Battle Royale the choice is the specialist: Researcher, Skeptic, Comedian, Optimizer. The UI shows the probabilities while it picks. No API key? It falls back to a backup referee.

Watch the Agent Battle Royale demo

PS: Do not put the TypeSafe key in the browser. Server side only.

Use cases and limitations

Where it fits

  • Routing: which agent, which team, which tool
  • Decision layer: approve, reject, escalate, roll back
  • Scoring: urgency, severity, quality of a draft
  • Guardrails: is this request safe enough to continue

Where it does not

  • Writing the email, the post, or the code
  • Open-ended chat
  • A problem where you cannot list the options
  • Anything you want to ship with zero human check on a high-stakes call

Jev is a smart if-statement, not a coworker. Your code still owns the action.

Also: garbage state, garbage judgment. If your incident notes are empty, the 81% rollback is theatre.

Alternatives

In last few weeks, we got a few alternatives of Jev:

  • OpenAI Decisions API. Announced at DevDay 2026. Same idea: finite answers, a probability, fast. Still a limited preview when I wrote this. Price and full schema were not as public as Jev's.
  • Laya. Open, you can run it yourself. Same shape of questions: choice, score, noul.
  • Kev. Open models that speak the same System One API, so you can point a Jev-style client at a local server.
  • Amazon Strands Decider 2B. AWS's open answer. Small model, runs local, picks from options you already listed. Code is on GitHub.

I started with hosted Jev because I wanted to feel the product, not rebuild the model. If you care about running local and not sending state out, look at Laya, Kev, or Strands Decider first.

What I would tell you to try

Do not wrap your whole agent in Jev on day one.

Pick one decision you already pay an LLM to make. The router. The "is this urgent" flag. The "which queue" call. Write the state. Lock three questions. Compare cost and whether your code got simpler.

If the answer is a paragraph, keep the LLM. If the answer is a label plus a number, try Jev.

Happy Learning!!