JEV Context Window Limits and Practical Implications
Two hidden context limits, not one, cause Jev pipelines to fail silently.

Jev's context window isn't one number, it's two, and mixing them up is why some pipelines compact gracefully while others fail without warning. A 32,000-token limit governs state plus the single longest query, while a separate 64,000-token figure caps state plus every query combined, and knowing which one binds a given workload is the difference between a system that degrades safely and one that just stops.
What Jev returns
Jev doesn't write sentences. It returns one of three typed outputs: a yes/no probability called noul, a category pick from a fixed list called choice, or a numeric score on a defined scale, and every one comes back as JSON, never as prose. That's the entire vocabulary. No paragraphs, no hedged explanations, no "it depends" tucked into a caveat clause.
It was built by TypeSafe AI, co-founded by Diogo Almeida, Erik Gafni, and Sasha Sheng, who co-invented RLHF and InstructGPT. TypeSafe came out of stealth on September 15, 2026, with $40 million in seed funding led by DCVC.
What Jev doesn't do matters just as much as what it does: no reasoning chains, no summarization, no theme extraction, no conversational back-and-forth. That absence is the reason its context window behaves so differently from a chat model's, and it's the throughline for everything that follows. TypeSafe claims Jev is 20–200x faster and 40–400x cheaper than frontier LLMs on decision-shaped work, with input tokens priced at $0.042 per million, output tokens free, and end-to-end latency of 70–500ms jev-router GitHub issue MindStudio.
The two-limit architecture: what the 32k and 64k numbers each govern
TypeSafe's own jaggedness documentation for jev-1.13 lays out two distinct ceilings, and conflating them is the single most common mistake in any Jev integration. The first is 32,000 tokens for state plus the single longest question in the request. The second is 64,000 tokens for state plus every question added together, the full per-request budget.
State gets counted once per request. For any workload with a growing or stateful exchange, the 32k figure is the operative ceiling. The 64k number only becomes relevant for a single large one-shot call carrying many sizable questions at once.
The confusion around these two numbers isn't accidental, it's structural. Platform listings routinely collapse them into one figure. TypeSafe's models page cites both, 64k per request and 32k for state plus the longest query, but OpenRouter lists just 32k, and Cloudflare's Jev documentation lists 32,000 tokens on its own. Different audiences land on different pages and walk away with different numbers in their heads. Neither number is wrong; they're just answering different questions. Which one binds depends entirely on the shape of the workload sitting on top of it, and that's the question the next section exists to answer.
How the 32k limit was measured and confirmed
Documentation is one thing. Live behavior is another, and this one was tested directly. On jev-1.13.0, a request carrying 32,878 input tokens went through cleanly; the next size step up was refused outright with an HTTP 400 error carrying the code max_tokens_exceeded. That result appears in pydantic/genai-prices PR #720, merged September 25, 2026, and it was confirmed independently the same day by pydantic/pydantic-ai PR #8740.
What matters here isn't the exact token count so much as the shape of the failure. Jev doesn't truncate quietly. It doesn't get vaguer or less confident as it approaches the edge. It throws an error and the call fails, in full, with nothing returned. Compare that to how a general-purpose LLM behaves under the same pressure: feed a large model too much input and its output quality slips, its answers wander, its grounding gets thinner. Feed Jev too much input and the request simply doesn't complete. Those are two different failure modes, and they call for two different engineering responses. A soft degradation can be caught downstream by a quality check. A hard cliff has to be caught before the request ever leaves the pipeline.
The 32,878-token figure is a useful calibration point to have in a back pocket during load testing.
The Pydantic AI bug: what happens when a framework believes the wrong number
That lesson got tested the hard way inside Pydantic AI. Before a fix landed, TypeSafeModel's profile in the framework set context_window=64_000, and the library's compaction logic keyed directly off that value. ctx.context_window_used read approximately 0.51 at the exact moment requests started failing, so the framework believed the model was roughly half full when it had already hit its hard ceiling.
The processor responsible for compacting a conversation once its window fills is built to trigger above 0.8 utilization. It never fired, because the utilization figure the framework was computing never crossed that threshold before the errors started arriving. This wasn't a synthetic benchmark either. A real support chat, running in production, failed at turn 29.
The fix, landed in pydantic/pydantic-ai PR #8740 and merged September 25, 2026, removes the hardcoded value entirely and pulls context_window: 32000 from genai-prices instead, a value introduced in the companion PR #720 from the same date and applied across jev-1.13.0, jev-latest, and jev-preview jev-router GitHub issue. Once the correct window was in place, the identical support chat compacted itself at turns 23 and 46, and ran the full 60-turn conversation to completion jev-router GitHub issue. Same conversation, same model, same content. The only variable that changed was which ceiling the framework knew to be true.
The lesson generalizes well past this one bug. Any framework, SDK, or integration that inherits or hardcodes 64k as Jev's context window is going to skip compaction silently and let conversations run straight into hard failures. It's a class of bug, and it will resurface anywhere a developer assumes the bigger of the two published numbers is the one that governs their workload.
How Jev's 32k ceiling sits relative to other models in the same workload category
Set against the current landscape of large-context models, Jev's ceiling looks small, deliberately so. Mainstream flagship models, GPT-5.5, Gemini 3.1 Pro, Claude Opus 4.8, and Claude Sonnet 5 among them, process on the order of a million tokens. Smaller, cheaper models tend to run somewhere between 128k and 256k. Meta's Llama 4 Scout advertises the largest window found in this research, at 10 million tokens.
Jev's 64k total budget is already well below any of those figures, and the operative 32k ceiling that governs most real conversations is smaller still. That's a design decision that follows directly from what Jev is built to do: it isn't meant to read a novel or reason across a sprawling document, it's meant to make a fast, narrow, typed decision. The state a builder sends into Jev has to be a deliberate reduction of the source material. An ordinary application DOM will blow past 32k tokens without much effort, and the reduction step that keeps it under that ceiling becomes part of the system, with its own bugs, its own edge cases, its own maintenance burden.
A second-order risk emerges here too. The jev-router integration, used to route requests across models, can itself shrink the effective context window available to whatever model sits underneath it. It's a routing-layer side effect, not a Jev limitation exactly, but it lands on the same pipeline and produces the same symptom: less usable context than the number on the label suggested.
None of this is a knock against Jev. The comparison exists to calibrate how much state discipline a given engineering task actually requires: Jev isn't a drop-in substitute for a large-context task, but a specialized instrument that demands state discipline its million-token counterparts simply don't require of a builder. A documented jev-router issue, filed September 20, 2026 as GitHub gargpratyush/jev-router issue #35, shows context dropping from 1M to 200k under jev-claude, with MCP tools consuming 18.8k tokens at startup versus 651 tokens under native Claude, with the cause listed as unknown jev-router GitHub issue jev-router GitHub issue.
Context rot: accuracy degradation before the hard limit is reached
Staying under 32,000 tokens doesn't guarantee good answers. TypeSafe's own materials name a phenomenon they call context rot: as state fills up with material the question doesn't actually need, accuracy drops, and TypeSafe lists this explicitly as a known weakness in jev-1.13. The same jaggedness documentation groups irrelevant context alongside other documented soft spots: literal interpretation, numerical precision, indirection, conflicting criteria, adversarial input.
Treating the 32k ceiling as a budget to spend down is the wrong mental model. TypeSafe's own framing is that irrelevant length actively lowers accuracy, making the ceiling a boundary to stay well clear of rather than a quota to fill. This isn't unique to Jev, either. The NoLiMa benchmark, run by researchers at LMU Munich and Adobe Research and presented at ICML 2025, found that once models couldn't lean on surface-level pattern matching, 11 of 13 tested LLMs dropped below half their baseline accuracy at just 32,000 tokens of context, with GPT-4o falling from 99.3% down to 69.7%. Long-context degradation is a general phenomenon across the field. Jev's hard stop just makes the consequence sharper and faster to hit, because there's no soft fade to mask the problem before it arrives.
The practical upshot: the 32k figure is the outer boundary, but the quality budget that actually matters is tighter than that number suggests. Filtering belongs before the state gets sent, not as cleanup afterward.
How question decomposition changes accuracy more than context size does
An independent evaluation published September 17, 2026, and covered at beri.net, tested this directly across 2,000 emails split between phishing and legitimate samples BERI independent eval MindStudio.
Then the evaluators changed the question, not the model. They split the single broad question into five atomic questions, fit weights on half the labelled data (1,000 emails), and tested the result on the held-out half BERI independent eval. Same 2,000 emails BERI independent eval MindStudio. Same model. The only thing that changed was how the question was structured going in.
That result says something structural about the whole category of decision, not just this one dataset. Jev underperforms a general-purpose model when handed a vague, sprawling question, and it outperforms that same model once the question is decomposed into narrow, atomic parts. Context window design, meaning what actually goes into state and how tightly the question gets scoped, ends up mattering more than raw speed or cost savings in determining whether Jev is even the right tool for a job. Put differently, the 32k ceiling forces a kind of discipline on question design, and when that discipline gets applied properly, it improves accuracy rather than just constraining it. The constraint, in this narrow sense, does some of the work for you. The 95.0% figure is specific to decomposed phishing classification with a fitted logistic regression, and the sourced caveat about statistical significance, per the BERI independent evaluation, should be retained.
Engineering patterns that keep workloads inside the limits
Documents larger than 32k tokens need to be chunked in application code before they ever reach Jev, because Jev was never built to ingest large documents whole. The segmentation or retrieval step belongs in the surrounding pipeline, not inside the model. For scale, the 32k state budget works out to roughly the equivalent of more than 500 posts at X's 280-character limit, which means the ceiling mostly bites on long documents, transcripts, and PDFs rather than on short-form text. TypeSafe's own guidance follows from that: send the post and the platform it came from. Filter aggressively before dispatch, not after.
On the Pydantic AI side, post-fix, the recommended pattern is to run the compact-when-the-window-fills processor behind a FallbackModel, with max_tokens set to leave real headroom under 32k rather than skimming right up against it. Jev's context window is not a single cap but two interlocking constraints (a 32k state-plus-longest-question limit and a 64k total budget), and which one binds a given workload determines whether a pipeline compacts correctly, fails silently, or runs all the way through. Once a conversation crosses the threshold, every subsequent Jev turn fails; behind a FallbackModel, those turns get routed to the language model instead, until the history gets compacted back down. And SummarizingCompaction needs an actual language model to do the summarizing, since Jev, by design, cannot produce summary text of its own.
Using Jev to compress large LLM conversations before they hit their own context ceilings is a meta-use here. That's Jev's low latency being used to gate what enters a different model's context, not to manage its own.
That kind of recursive use carries its own risk, though. TypeSafe warns that accuracy drops when state carries a large volume of irrelevant detail, and a compactor that asks Jev to inspect a big, noisy transcript can degrade its own judgment in the process. Pre-grouping candidates, fitting the state carefully, and testing behavior as context grows all matter here just as much as they do in the primary use case. The harness anchors on reported usage from the last request and estimates newer content at roughly 4 chars/token, which undercounts Jev's JSON state by about 28%, so the trigger should fire earlier than the measured limit to absorb that undercount. A third-party pattern uses Jev as a compactor to compress large LLM conversations before they hit context limits, with one implementation reported taking a Claude session from nearly 1M tokens to 86k tokens in roughly 1 second. The jev-skill-suggestion mod, by Daniel Avila on Medium, removes the full skill listing from Claude Code prompts, using a Jev choice question to pick at most one skill per prompt in approximately 330ms so that only that skill's SKILL.md reaches the context window Medium / Daniel Avila.
Versioning and platform listing risks that reintroduce the same bugs
jev-latest currently resolves to jev-1.13.0. When a new version ships, every threshold that's been calibrated against the current model moves along with it, and nothing in an application's own logs will flag that the change happened. The response object's model field reports the specific versioned ID that answered a given request, and that field is worth logging on every call, with evaluations replayed before any upgrade goes live. Pinning to jev-1.13.0 is the safer default for anything calibrated or consequential.
The platform listing divergence noted earlier compounds this risk rather than resolving it. TypeSafe's own models page states both figures, 64k per request and 32k for state plus the longest query. OpenRouter lists 32k on its own. Failproof AI, as verified September 22, 2026, states both limits together. Any integration that pulls its context window value from a platform listing, rather than from genai-prices directly (the context_window: 32000 value has been present since version 0.1.9), risks working from the wrong number entirely. Verifying the source of truth before designing compaction logic isn't optional caution, it's the step that bug demonstrated was missing.
Two further caveats belong in any production plan. Published direct-API rate limits are 250,000 tokens per second and 1,200 requests per minute, but TypeSafe states that these figures are adjusting dynamically, so they should be treated as a dated snapshot rather than a fixed contract, with 429 responses handled through standard SDK retry logic. Jev remains available through OpenRouter and the Vercel AI Gateway as an alternate path in, but any pipeline design should verify current access status rather than assume the onboarding route that worked last week still works today.
Sources
- Take Jev's `context_window` from genai-prices and document compacting a Jev conversation by DouweM · Pull Request #8740 · pydantic/pydantic-ai
- Add TypeSafe Jev's 32k `context_window` by DouweM · Pull Request #720 · pydantic/genai-prices
- 12 Jev Use Cases Tested: Where This Decision-Only AI Actually Fits | MindStudio
- Context window drops from 1M to 200k under jev-claude · Issue #35 · gargpratyush/jev-router
- Jev Skill Suggestion keeps your Claude Code skills out of context until one is needed | by Daniel Avila | Sep, 2026 | Medium
- Jev pricing, context window and rate limits · Failproof AI
- TypeSafe (Jev) | Pydantic Docs


