JEV report
JEV ExplainedLong read

JEV Reasoning Modes and When to Use Each

Three decision modes optimized for different types of judgment calls.

Contributing Editor · · 11 min read
Cover illustration for “JEV Reasoning Modes and When to Use Each”
JEV Explained · October 2, 2026 · 11 min read · 2,430 words

Jev is a model that answers typed yes/no, multiple-choice, and rating questions about a piece of state, instead of generating text. Its three output modes, Choice, Score, and Noul, are not interchangeable: each returns a structurally different signal suited to a different class of decision, and picking the wrong one costs accuracy, speed, or both.

What Jev is and why it exists as a separate layer

An agent loop that calls a large language model to select a tool, flag risk, judge a result, and decide when to stop is spending generative cost on decisions that are, structurally, bounded. The model is asked to pick from a small set of possibilities, yet it still produces its answer one token at a time, as if the answer were an essay rather than a label. The alternatives have real costs of their own. Hand-written rules are cheap to run but break as the language users send in shifts; traditional classifiers need labeled training data and periodic retraining; even constrained to JSON mode, generation is still produced token by token under the hood.

TypeSafe AI built Jev to sit in the gap those options leave open. Released in September 2026, Jev is pitched not as a substitute for large language models but as a narrower layer purpose-built for fast, structured decisions, with tool calls and routing steps inside agent loops as the primary use case. The company's founder, Diogo Almeida, helped invent reinforcement learning from human feedback at OpenAI, the training method behind InstructGPT and ChatGPT, and TypeSafe AI came out of stealth in mid-September 2026, on September 15, with Jev as its first released model.

TypeSafe calls Jev a "System One model," borrowing a distinction between fast, intuitive judgment and slow, deliberate reasoning: Jev never generates text, it answers typed questions about a given state. The name Jev is also a nod to William Stanley Jevons, whose paradox holds that making something cheaper tends to increase how much of it gets used in total, and the bet behind the model is that decisions currently routed to expensive generative models will get made far more often once the cost of making them drops.

The mechanism that makes this different from existing approaches is architectural. LLM JSON mode still runs a generative model that produces tokens, with the schema applied afterward as a filter or a retry mechanism when the output doesn't parse. Jev's training method, which TypeSafe calls RLCD, or Reinforcement Learning for Calibrated Decisions, is described by the company as a descendant of RLHF adapted to decision-space outputs rather than text-space ones, optimizing the model solely for calibration rather than for fluent generation. RLCD has not been documented in a published paper or technical report, so the calibration claims attached to Jev currently rest on TypeSafe's own evaluation rather than on independently reproduced results. That distinction matters for how much weight a given confidence number should carry, a point the rest of this piece returns to directly.

The practical consequence is an interface built around a constraint. Unstructured state goes in: a string, a JSON object, an array of text. Typed, probabilistic decisions come out, with the output shape fixed before the model runs rather than checked and corrected after the fact.

The Three Output Primitives

Choice, Score, and Noul look, at first glance, like three ways of asking the same question, but they are not. Each returns a structurally different signal, and each signal is useful for a structurally different class of problem.

Choice selects one option from a predefined set and returns a probability for every option in that set, along with a single confidence score. The winning label is only part of the answer. The full probability distribution is the signal that actually matters, because a narrow margin between two options describes a different situation entirely than a commanding one, even when both cases technically produce the same winning label. A real Choice response shows this concretely: {"choice": "billing", "probabilities": {"billing": 0.52, "technical": 0.46, "sales": 0.02}, "confidence": 0.18}. Billing wins, but a confidence of 0.18 means automatic routing on that result alone would be reckless, regardless of which label came out on top.

Score places an input on an ordered rubric that the developer defines, something like Calm, Frustrated, and Very Angry, and returns a scalar score, a legend mapping that score to its labels, and a probability distribution across the rubric's levels. A sample response from the Jeeves quickstart shows the full shape: {"score": 1.5, "legend": {"0": "Calm", "1": "Frustrated", "2": "Very angry"}, "probabilities": {"0": 0.04, "1": 0.43, "2": 0.54}, "confidence": 0.75}. This output is ordinal rather than categorical: a score of 1.5 on the 0–2 scale communicates something a pure label cannot, namely that the input sits between levels.

Noul, TypeSafe's name for its Boolean decision type, returns a single probability between 0 and 1 that a stated proposition is true. There is no competing label to weigh against another, only a calibrated probability that a given statement holds, with the threshold for acting on that probability left entirely to the developer. For the proposition "does this need urgent human attention," a Noul response might read simply {"noul": 0.72}, a number that feeds directly into a branch condition in application code. All three question types can be sent in the same request and evaluated in parallel, as long as none of the answers depends on another; any decision where one answer determines the next question belongs in code, not inside Jev.

Knowing the output shape of each mode is necessary, but it does not guarantee a correct decision. A Choice response can return department=billing with high stated confidence when the correct routing was actually technical support. The schema guarantees that the output is well-formed. It says nothing about whether the underlying judgment was right. That gap, between type-valid output and correct output, is the failure mode the next section builds on directly.

Matching each mode to its decision class

The selection rule falls directly out of what each mode returns: use Choice for picking one option out of a set, use Score for an ordered assessment of severity, priority, or quality, and use Noul for judging whether a single proposition is true.

Choice fits any decision that is genuinely a selection among mutually exclusive categories. A support ticket needs to land in exactly one queue: returns, shipping, or billing. A citation check needs a verdict on whether a quoted passage supports a claim, contradicts it, or does neither. A routing layer needs to decide which model tier, fast or powerful, should handle an incoming request. It breaks down the moment they don't: if billing and technical support overlap in practice, Jev still has to pick one, and the resulting probability distribution will show the ambiguity as a near-even split rather than a clean winner.

Score fits decisions where the dimension being measured is ordered and the distance between levels carries information. Reranking a passage for relevance to a query calls for a rating on a defined scale rather than a binary accept-or-reject. Scoring a resume or a claim against a rubric works the same way, producing a structured evaluation rather than a single cutoff. Frustration or urgency triage is a particularly good fit, because the scalar output lets downstream code apply a graduated response rather than forcing every case through a binary branch. Score breaks down when the rubric's levels aren't well-ordered, or when what's actually needed is a single yes-or-no gate rather than a gradient; in that case, Noul is the correct tool.

Noul fits any decision that reduces to the truth of one proposition. A tool-call guardrail can ask whether a proposed shell command matches the agent's stated task. A risk classifier can ask whether a given action is irreversible, whether it falls off-task, or whether it mutates state. A routing gate can ask whether a given prompt needs a larger model. Noul breaks down when a question is actually a compound condition in disguise; that kind of logic belongs in code, or should be split into multiple separate Noul questions rather than forced into one.

The modes are not mutually exclusive at the request level. Choice and Score can run side by side in a single call: a support triage system might ask which queue should own a ticket (Choice) and how frustrated the customer sounds (Score) simultaneously, since the answers are independent so parallel evaluation applies.

One category of operation belongs in neither mode, and in no part of Jev at all. Arithmetic, counting, date comparison, and exact string matching are deterministic operations, and none of them should be delegated to any model, regardless of which primitive is in use. TypeSafe's own documentation makes this same recommendation: keep computationally exact operations in code, and reserve Jev for narrowly scoped semantic judgments. The resulting layering is the architectural principle the rest of an agent system should be built around: deterministic rules belong in code, bounded semantic judgment belongs in Jev, open-ended reasoning and communication belong in language models, and independent authorization belongs around any action with real consequences.

Confidence scores as control flow

Diagram: Reading a Confidence Score: When to Act, Escalate, or Queue. Visualizes: Illustrate a three-zone decision ladder built around the confidence signal Jev returns.

Picking the right mode is only half the job. The probability distribution each mode returns, not just the winning label, is what makes automated routing safe at scale. Code that reads only the top label and discards the distribution around it throws away the exact signal that tells it when not to act.

In the middle band, the right move is to escalate, either to a slower and more capable model or to a request for human confirmation. At low confidence, the case belongs in a queue for human review, or the system should gather more information before deciding at all. Where exactly those thresholds sit is a decision for the developer, and it should track the consequence of acting wrongly: a dashboard label can tolerate a weaker signal than a command that deletes data or charges a customer's card.

TypeSafe's documentation illustrates the pattern with a short branch structure:

if urgent > 0.9 and owner == "engineering":
 page_on_call()
elif confidence < 0.5:
 add_to_queue(owner)

A high urgency score tied to a specific team triggers an immediate page; a low-confidence result, regardless of what it says, goes to a human queue instead of being acted on automatically.

Calibration is the operative concept here: a model is calibrated when its stated confidence actually tracks its observed accuracy. None of this works unless the confidence numbers themselves are trustworthy. That property should be measured on a system's own traffic rather than assumed from a vendor's published benchmarks, and it needs to be checked separately for each decision type, then rechecked after any change to the model, the questions asked, or the distribution of inputs coming in.

A recent study gives this pattern a concrete, published instantiation. The jev-as-a-Judge paper, published by researchers at Carnegie Mellon in September 2026, builds a frozen cascade that accepts a confident verdict as final and escalates only the uncertain ones, retaining close to full accuracy at a fraction of the cost of sending every single item to a frontier judge model. The confidence signal is precisely what makes that selective escalation work. Across several of the benchmarks in that study, the accuracy gap between Jev and the strongest LLM comparator concentrates almost entirely in the low-confidence decisions, which is exactly the set the cascade is designed to catch and route elsewhere.

The same tiered logic appears in agent security tooling built around Jev. A tool called pi-warden sends an agent's task, its stated plan, and a pending command to Jev as four separate typed Noul questions: is this action irreversible, is it off-task, does it mutate anything, and what is its scope, all evaluated before a bash, write, or edit command is allowed to execute. A confident, read-only answer lets the action proceed; a destructive or uncertain one pauses the agent for human approval. LangChain has built a comparable pattern into its own middleware, where an external Jev-based judge inspects a proposed tool call before it runs, and Pydantic's own guidance warns that thresholds for this kind of gate need to be set from labeled examples rather than chosen arbitrarily.

The think/no-think dimension that cuts across all three modes

Mode selection and confidence thresholds answer what Jev decides and how that decision gets acted on. A separate question is whether Jev reasons through a chain of intermediate steps before producing an answer. That choice trades latency against accuracy on inputs that are hard or unfamiliar, and the correct default is not simply "always think".

Research on this tradeoff, published as the NoThinking study (arXiv 2505.13417), found that skipping the reasoning chain entirely produces accuracy comparable to a full reasoning pass on relatively simple problems, and slightly outperforms it on the very easiest ones, while generating responses that are significantly shorter. For the kind of narrow-scope task Jev is built for, routing, classification, guardrail checks, no-think may well be the right default behavior, with full reasoning reserved for decisions that are genuinely ambiguous or fall outside the model's training distribution. TypeSafe's Jev API does not expose the think/no-think toggle directly; that choice is made internally by the model.

PostHog's open-source Jeeves, released in late September 2026, takes the opposite position and exposes the tradeoff as a setting the developer controls directly, something Jev's closed API does not offer. Jeeves trains a Qwen3.5 model, using LoRA and a pointer head, with a method called CISPO that teaches it to reason before answering. Without thinking enabled, it returns a result in well under a second; with full thinking turned on, producing over a thousand reasoning tokens on average, accuracy improves but median latency climbs to 3.3 seconds, with the slowest cases, the p90, stretching considerably further. The project exposes max_think and nothink_threshold parameters specifically so that developers can tune this tradeoff for systems where latency is the binding constraint.

The resulting performance picture is mixed rather than one-sided. Jeeves scores 0.935 on the 231-item JevBench public tier and 0.889 on out-of-domain test data, beating Jev's 0.866 and 0.857; on rule-structure and contrastive-policy tasks Jeeves hits 1.000 against Jev's 0.885 and 0.963. Knowledge questions still trail Jev: Jeeves scores 0.793 against Jev's 0.900, suggesting Jeeves holds an edge on structured rule-following while Jev retains an advantage on factual recall. Developers choosing between an exposed reasoning toggle and a closed, internally managed one are, in effect, choosing how much control they want over that latency-accuracy boundary, and how much they're willing to trust a vendor's internal defaults to draw it correctly on their behalf.

Sources

  1. GitHub - PostHog/jeeves: Jeeves
  2. Jev, Clearly Explained
  3. Jev vs. LLMs: When AI Moves from Generation to Decision-Making
  4. Jev AI Security: Why Decision Models Could Change Agent Security
  5. JEV-as-a-Judge: Accept When Confident, Escalate When Unsure
  6. AdaptThink: Reasoning Models Can Learn When to Think
Filed underJEV Explained

More in JEV Explained