← All AI resources

A new model category · Introducing System One Models & Jev

Jev is a new kind of model for decisions inside software

It does not chat, write code, or generate prose. Jev reads state, answers bounded questions, and gives your application typed decisions with probabilities.

Published by Learnetto team · Launched September 15, 2026

TypeSafe AI is betting that many production AI calls do not need another general-purpose LLM. They need a fast judgment that code can inspect, combine, threshold, and reject. Jev is built for that narrower job.

Model
jev-latest
Output
Choice, Score, or Noul
Interface
POST /v1/systemone
Launch status
Early access
Vercel adoption
Nearly 13% of paid teams in 24 hours

Why Jev matters

Most LLM-based automation asks a text generator to make a small decision, forces it into JSON, validates the result, and discards the prose. Jev removes that detour. The answer space is defined before the call, and the model returns probabilities your code can use directly.

That makes Jev more like a semantic decision primitive than a chatbot. Your code still owns control flow, permissions, calculations, and side effects. Jev handles the part ordinary rules struggle with: interpreting messy language or application state.

How Jev works

A request contains state plus one or more typed questions. Jev evaluates the questions independently against the same state and returns all answers in one response.

Use Choice when one option must win, Score for a position on an ordered rubric, and Noul for the probability that a condition is true. Code can then act above a tested threshold, ask for review, or route uncertain cases elsewhere.

Where Jev fits

  • Route support tickets, leads, documents, or agent requests to a known handler.
  • Choose a model or tool for each request while keeping the probability distribution for auditability.
  • Score risk, relevance, urgency, quality, or policy fit before code takes an action.
  • Gate agent tool calls and send uncertain or high-risk actions to a person.
  • Select a source value or candidate in code instead of asking a model to invent one.

What changed after launch

Vercel reported that Jev became the fastest-adopted model in AI Gateway history. Within its first 24 hours on the gateway, nearly 13% of paid teams had used it, more than twice the share of any previous model launch on that platform. That is platform-reported adoption, not proof that those teams kept Jev in production or improved business outcomes.

The early experiments point to a consistent division of labor: use Jev for frequent bounded choices, scores, and gates, while code keeps permissions and side effects and a generative model handles prose or longer reasoning. The pattern is promising because it is measurable, but every threshold still needs task-specific evaluation.

What independent testing found

A preregistered community pilot run on September 17, 2026 compared hosted Jev 1.13 with local GLiNER2.5 across three 100-example classification slices. Jev reached 91% accuracy on AG News and 87% on the 72-label Banking77 slice, ahead of GLiNER's 70% and 61% in that setup.

The result was mixed rather than universal. Jev reached 48% accuracy on emotion classification versus GLiNER's 44%, a difference the study could not distinguish statistically. Jev's probability quality was worse there: it assigned zero probability to the true label on 16% of examples. The comparison also used a hosted service from France against local Apple CPU inference, so its latency figures are not a hardware-normalized benchmark.

What to be cautious about

Typed output prevents schema mistakes, not wrong judgments. TypeSafe's speed, cost, and quality comparisons are vendor-published results for System One-shaped workflows. The independent pilot supports Jev on some classification tasks and exposes a calibration failure on another, so neither source proves broad superiority.

TypeSafe's Jev 1.13 limitations page says Jev is not a calculator, can read conditions too literally, struggles with date and time comparisons, and can lose accuracy with unrelated context. It accepts text rather than images, audio, or video, and does not generate explanations, prose, code, or long reasoning traces. Keep exact math and control flow in code, minimize the state, and test probabilities and thresholds on representative labeled data before automating consequential decisions.

What people are already building with Jev

These examples come from public posts on X. They show the range of early experiments, not independently audited production results. Matt Van Horn's roundup helped surface several of the original posts; Learnetto checked each post directly and wrote its own descriptions.

Example on X · Hrishi Mittal

Citation verification and content intelligence for AI search

Hrishi is testing Jev inside Evatype for citation verification, search-keyword classification, content validation, and market inference, with early results suggesting it could replace slower LLM calls in these bounded workflows.

Example on X · Kyle Jeong

Browser actions for fractions of a cent

A Stagehand browser loop sends the accessibility tree and possible actions to Jev, which selects the next action before Stagehand executes it. The posted task cost $0.001.

Example on X · Jarrod Watts

A trading bot making 300 ms decisions

Jev reads an asset-pair price feed, chooses buy or sell, and passes that bounded decision to code placing real trades on an on-chain order book. This is a demo, not evidence that the strategy is safe or profitable.

Example on X · Ryan Vogel

Classifying 1,500 personal emails

A high-volume email test uses Jev to classify 1,500 real messages, showing where typed labels can be more useful than generated prose.

Example on X · Raihan Khan

DiffJury reviews whether a pull request is safe to merge

DiffJury accepts a public pull-request link and uses Jev to judge whether it looks safe to merge or should receive human review.

Example on X · Kun Chen

Production routing between agent configurations

FirstMate replaced an LLM-based dispatch step with Jev for choosing an agent harness, model, and reasoning effort. Its creator reports matching the previous answers while sharply reducing wall time and cost.

Example on X · Dev Agrawal

A staged code-review dashboard

Jev Review screens Git diffs, follows strong signals through staged judgments, and presents the result in a local dashboard.

Example on X · Joey Kudish

Evidence verification inside an agent harness

A proof-of-concept MCP uses Jev to verify claims against evidence, screen material before it enters context, and rank candidates by meaning.

Example on X · Niaz Morshed

A score-and-improve feedback loop for coding agents

A local-first MCP plugin lets coding agents ask Jev for quality scores, improve their work, and repeat the loop across multiple review dimensions.

Example on X · Vishesh Baghel

A personalized Hacker News ranking

Jev scores stories for technical depth, drama, practical utility, AI slop, novelty, and career relevance so readers can rerank the front page with sliders.

Example on X · Aaron Levin

Cross-platform computer use

A computer-use prototype feeds on-screen text regions to Jev to choose actions across operating systems. Its creator reports much lower latency and cost than an Opus-based loop.

Example on X · Joogie

A Jev-powered Wikipedia racing agent

An agent-browser and Jev loop races from one Wikipedia topic to another while code constrains the links and actions Jev may choose.

Example on X · Marcus Lowe

Playing Tetris at real-time speed

A Tetris demo uses Jev's low-latency decisions quickly enough to control falling pieces in real time.

Example on X · Gregor Zunic

Finding flights in seven seconds for $0.0039

Browser Use built a small open-source agent where Jev selects from a fresh DOM-derived action space on every step and a small LLM is used only when text must be generated.

Example on X · Tamara Tran

Instant context compaction by scoring tool calls

Instead of asking an LLM to summarize a nearly full agent context, the plugin asks Jev to score tool calls and removes the ones judged irrelevant.

Example on X · Alex Volkov

Compressing a Claude session from nearly 1M to 86K tokens

Alex Volkov tested the Jev compaction plugin in Claude and reports that it reduced a nearly one-million-token session to 86,000 tokens in about one second.

Example on X · Guillermo Rauch

A production safety reviewer for shell commands

Vercel is testing Jev as the safety reviewer that evaluates commands before fx auto mode executes them. Guillermo Rauch reports an 18x p95 speed improvement and higher accuracy than the current reviewer.

Example on X · Sydney Runkle

Model routing and agent middleware with LangChain

LangChain shows how Jev can choose between models, gate risky tools, and make repeated decisions between larger generative-model calls while leaving probabilities in agent state.

Example on X · Kush Bhuwalka

Filtering irrelevant chunks after RAG retrieval

A simple RAG pattern retrieves chunks normally, then asks Jev whether each chunk is relevant and drops the ones that fail the threshold before generation.

Example on X · Wuyang Zhou

Fast reactions and slow planning in Minecraft

A Minecraft agent pairs Jev for immediate reactions with GPT-6 Astra for longer-horizon planning, including combat with multiple enemies.

Example on X · Nader Dabit

A predictive launcher that reranks on every keystroke

A launcher sends natural-language intent to Jev on every keystroke so a query such as 'the PDF I just downloaded' can place the right file first in roughly 100 ms.

Example on X · Tobi Lütke

Running a Jev-like decision model entirely in the browser

Reflex recreates the typed-decision interface with a local Qwen model on WebGPU, providing an offline example of the same bounded-decision pattern without a hosted API.

Example on X · Matt Van Horn

A community roundup of Jev experiments

A wider roundup covers browser agents, context compaction, safety review, model routing, and other early attempts to use Jev for repeated judgment calls.

Keep learning on Learnetto

Primary sources

TypeSafe AI launch, September 15, 2026

TypeSafe Jev 1.13 limitations, reviewed September 17, 2026

Vercel AI Gateway adoption report, verified September 27, 2026

Independent Jev versus GLiNER pilot, run September 17, 2026