A new model category · Introducing System One Models & Jev
Jev is a new kind of model for decisions inside software
It does not chat, write code, or generate prose. Jev reads state, answers bounded questions, and gives your application typed decisions with probabilities.
TypeSafe AI is betting that many production AI calls do not need another general-purpose LLM. They need a fast judgment that code can inspect, combine, threshold, and reject. Jev is built for that narrower job.
- Model
- jev-latest
- Output
- Choice, Score, or Noul
- Interface
- POST /v1/systemone
- Launch status
- Early access
- Vercel adoption
- Nearly 13% of paid teams in 24 hours
Why Jev matters
Most LLM-based automation asks a text generator to make a small decision, forces it into JSON, validates the result, and discards the prose. Jev removes that detour. The answer space is defined before the call, and the model returns probabilities your code can use directly.
That makes Jev more like a semantic decision primitive than a chatbot. Your code still owns control flow, permissions, calculations, and side effects. Jev handles the part ordinary rules struggle with: interpreting messy language or application state.
How Jev works
A request contains state plus one or more typed questions. Jev evaluates the questions independently against the same state and returns all answers in one response.
Use Choice when one option must win, Score for a position on an ordered rubric, and Noul for the probability that a condition is true. Code can then act above a tested threshold, ask for review, or route uncertain cases elsewhere.
Where Jev fits
- Route support tickets, leads, documents, or agent requests to a known handler.
- Choose a model or tool for each request while keeping the probability distribution for auditability.
- Score risk, relevance, urgency, quality, or policy fit before code takes an action.
- Gate agent tool calls and send uncertain or high-risk actions to a person.
- Select a source value or candidate in code instead of asking a model to invent one.
What changed after launch
Vercel reported that Jev became the fastest-adopted model in AI Gateway history. Within its first 24 hours on the gateway, nearly 13% of paid teams had used it, more than twice the share of any previous model launch on that platform. That is platform-reported adoption, not proof that those teams kept Jev in production or improved business outcomes.
The early experiments point to a consistent division of labor: use Jev for frequent bounded choices, scores, and gates, while code keeps permissions and side effects and a generative model handles prose or longer reasoning. The pattern is promising because it is measurable, but every threshold still needs task-specific evaluation.
What independent testing found
A preregistered community pilot run on September 17, 2026 compared hosted Jev 1.13 with local GLiNER2.5 across three 100-example classification slices. Jev reached 91% accuracy on AG News and 87% on the 72-label Banking77 slice, ahead of GLiNER's 70% and 61% in that setup.
The result was mixed rather than universal. Jev reached 48% accuracy on emotion classification versus GLiNER's 44%, a difference the study could not distinguish statistically. Jev's probability quality was worse there: it assigned zero probability to the true label on 16% of examples. The comparison also used a hosted service from France against local Apple CPU inference, so its latency figures are not a hardware-normalized benchmark.
What to be cautious about
Typed output prevents schema mistakes, not wrong judgments. TypeSafe's speed, cost, and quality comparisons are vendor-published results for System One-shaped workflows. The independent pilot supports Jev on some classification tasks and exposes a calibration failure on another, so neither source proves broad superiority.
TypeSafe's Jev 1.13 limitations page says Jev is not a calculator, can read conditions too literally, struggles with date and time comparisons, and can lose accuracy with unrelated context. It accepts text rather than images, audio, or video, and does not generate explanations, prose, code, or long reasoning traces. Keep exact math and control flow in code, minimize the state, and test probabilities and thresholds on representative labeled data before automating consequential decisions.
What people are already building with Jev
These examples come from public posts on X. They show the range of early experiments, not independently audited production results. Matt Van Horn's roundup helped surface several of the original posts; Learnetto checked each post directly and wrote its own descriptions.
Citation verification and content intelligence for AI search
Hrishi is testing Jev inside Evatype for citation verification, search-keyword classification, content validation, and market inference, with early results suggesting it could replace slower LLM calls in these bounded workflows.
View Hrishi Mittal's post on X
Browser actions for fractions of a cent
A Stagehand browser loop sends the accessibility tree and possible actions to Jev, which selects the next action before Stagehand executes it. The posted task cost $0.001.
View Kyle Jeong's post on X
A trading bot making 300 ms decisions
Jev reads an asset-pair price feed, chooses buy or sell, and passes that bounded decision to code placing real trades on an on-chain order book. This is a demo, not evidence that the strategy is safe or profitable.
View Jarrod Watts's post on X
Classifying 1,500 personal emails
A high-volume email test uses Jev to classify 1,500 real messages, showing where typed labels can be more useful than generated prose.
View Ryan Vogel's post on X
DiffJury reviews whether a pull request is safe to merge
DiffJury accepts a public pull-request link and uses Jev to judge whether it looks safe to merge or should receive human review.
View Raihan Khan's post on X
Production routing between agent configurations
FirstMate replaced an LLM-based dispatch step with Jev for choosing an agent harness, model, and reasoning effort. Its creator reports matching the previous answers while sharply reducing wall time and cost.
View Kun Chen's post on X
A staged code-review dashboard
Jev Review screens Git diffs, follows strong signals through staged judgments, and presents the result in a local dashboard.
View Dev Agrawal's post on X
Evidence verification inside an agent harness
A proof-of-concept MCP uses Jev to verify claims against evidence, screen material before it enters context, and rank candidates by meaning.
View Joey Kudish's post on X
A score-and-improve feedback loop for coding agents
A local-first MCP plugin lets coding agents ask Jev for quality scores, improve their work, and repeat the loop across multiple review dimensions.
View Niaz Morshed's post on X
A personalized Hacker News ranking
Jev scores stories for technical depth, drama, practical utility, AI slop, novelty, and career relevance so readers can rerank the front page with sliders.
View Vishesh Baghel's post on X
Cross-platform computer use
A computer-use prototype feeds on-screen text regions to Jev to choose actions across operating systems. Its creator reports much lower latency and cost than an Opus-based loop.
View Aaron Levin's post on X
A Jev-powered Wikipedia racing agent
An agent-browser and Jev loop races from one Wikipedia topic to another while code constrains the links and actions Jev may choose.
View Joogie's post on X
Playing Tetris at real-time speed
A Tetris demo uses Jev's low-latency decisions quickly enough to control falling pieces in real time.
View Marcus Lowe's post on X
Finding flights in seven seconds for $0.0039
Browser Use built a small open-source agent where Jev selects from a fresh DOM-derived action space on every step and a small LLM is used only when text must be generated.
View Gregor Zunic's post on X
Instant context compaction by scoring tool calls
Instead of asking an LLM to summarize a nearly full agent context, the plugin asks Jev to score tool calls and removes the ones judged irrelevant.
View Tamara Tran's post on X
Compressing a Claude session from nearly 1M to 86K tokens
Alex Volkov tested the Jev compaction plugin in Claude and reports that it reduced a nearly one-million-token session to 86,000 tokens in about one second.
View Alex Volkov's post on X
A production safety reviewer for shell commands
Vercel is testing Jev as the safety reviewer that evaluates commands before fx auto mode executes them. Guillermo Rauch reports an 18x p95 speed improvement and higher accuracy than the current reviewer.
View Guillermo Rauch's post on X
Model routing and agent middleware with LangChain
LangChain shows how Jev can choose between models, gate risky tools, and make repeated decisions between larger generative-model calls while leaving probabilities in agent state.
View Sydney Runkle's post on X
Filtering irrelevant chunks after RAG retrieval
A simple RAG pattern retrieves chunks normally, then asks Jev whether each chunk is relevant and drops the ones that fail the threshold before generation.
View Kush Bhuwalka's post on X
Fast reactions and slow planning in Minecraft
A Minecraft agent pairs Jev for immediate reactions with GPT-6 Astra for longer-horizon planning, including combat with multiple enemies.
View Wuyang Zhou's post on X
A predictive launcher that reranks on every keystroke
A launcher sends natural-language intent to Jev on every keystroke so a query such as 'the PDF I just downloaded' can place the right file first in roughly 100 ms.
View Nader Dabit's post on X
Running a Jev-like decision model entirely in the browser
Reflex recreates the typed-decision interface with a local Qwen model on WebGPU, providing an offline example of the same bounded-decision pattern without a hosted API.
View Tobi Lütke's post on X
A community roundup of Jev experiments
A wider roundup covers browser agents, context compaction, safety review, model routing, and other early attempts to use Jev for repeated judgment calls.
View Matt Van Horn's post on X
Keep learning on Learnetto
Primary sources
TypeSafe AI launch, September 15, 2026
TypeSafe Jev 1.13 limitations, reviewed September 17, 2026
Vercel AI Gateway adoption report, verified September 27, 2026