System One models vs LLMs: most AI calls should decide, not talk
Routing, gating, judging and choosing an agent's next step are decisions, not writing. Splitting decision from generation is the right architecture, and TypeSafe's Jev is the first product built for it. Trust will come from independently measured calibration, not a promise that it cannot hallucinate.
TL;DR
- A large share of the model calls inside production software are decisions: route this ticket, block this tool call, score this answer, pick the agent’s next step. Asking a text generator to write out those decisions is the wrong tool for the job.
- System One models, a category TypeSafe AI introduced with Jev on 15 September 2026, return a typed value with probabilities instead of text. On TypeSafe’s own workflow evals, Jev matched mid-tier LLMs on accuracy at a small fraction of the cost and time, and trailed the strongest ones.
- The architecture is right. The trust case is not yet made. Calibration is asserted by the vendor and measured by nobody else in public, and “can’t hallucinate” describes a schema guarantee, not correctness.
- What would settle it: independent reliability diagrams on real workloads, per-version calibration data, and a portable interface that lets buyers swap the decision model.
This is an opinion piece. The opinion: most of the AI calls that software makes should decide, not talk. When a program asks a model whether a refund request is valid, which queue a ticket belongs in, or whether an agent’s shell command is safe, the program needs a value it can branch on and an honest estimate of how likely that value is to be right. It does not need prose. For three years the default has been to send those questions to a large language model and parse whatever string comes back. Splitting the decision out of generation is the better architecture, and System One models are the first product built around that split. But the category has launched with the wrong headline promise, and it will not deserve trust until someone other than the vendor shows that its probabilities mean what they say.
What a System One model is
TypeSafe AI’s documentation defines System One models as “a class of AI models built to make fast, structured decisions that software can use directly.” A System One model reads a “state”, which is the text or JSON describing the situation, and returns “typed answers and probabilities.” Jev is the first one. TypeSafe released it in early access on 15 September 2026, according to the launch post by founder Diogo Almeida, and DCVC announced the same day that it led a $40 million seed round in the San Francisco company.
The developer fixes the answer space before the call through three primitives. A Choice picks one option from a list, a Score places the state on a defined scale, and a Noul returns the probability that a yes-or-no statement is true. The docs are explicit about what is given up: System One models “do not write replies, produce code, or generate explanations of their reasoning.”
The training method is the other half of the pitch. TypeSafe calls it reinforcement learning for calibrated decisions, or RLCD. Its AI primer says the goal is that “higher probability should correspond to a greater chance that the answer is correct”, and explains calibration with a simple rule: outcomes given a probability of 0.8 “should occur about 80% of the time.” The name borrows from Daniel Kahneman. In his 2002 Nobel lecture, Kahneman described the operations of System 1 as “fast, automatic, effortless, associative, and difficult to control or modify.” For a fuller walkthrough of the mechanics, see the System One models explainer; for the launch details, the news report on Jev.
Decisions dressed up as text
Look at where model calls sit in a typical production system and a pattern appears. A router decides which model or team handles a request. A guardrail decides whether an input or output is allowed. A judge decides whether an answer is good enough to ship. An agent harness decides which tool to call next, and a safety layer decides whether that call may run. Each of these is a classification with a small, known answer space. In most stacks today, each is implemented as a prompt to a generative model followed by a parser.
The evidence that these are decisions, not writing tasks, predates Jev. RouteLLM, a 2024 paper from researchers including Ion Stoica and Joseph Gonzalez, framed model routing as a learned choice between a strong and a weak model and reported that trained routers cut costs “by over 2 times in certain cases” without hurting response quality. The paper that popularized LLM-as-a-judge found GPT-4 judges reached “over 80% agreement” with humans, and also catalogued the “position, verbosity, and self-enhancement biases” that come from asking a text generator to render a verdict.
The early Jev reports cluster in exactly these slots. TechCrunch reported that Vercel had been using OpenAI’s Luna 5.6 “to run a classifier to review commands for safety”, and that when it switched to Jev, “it got results five to 18 times more quickly and with greater accuracy.” LangChain’s engineering post proposes using “an LLM for open-ended reasoning and generation, and Jev for fast, structured decisions along the way”, naming model routing and gating risky tool calls as the use cases.
The strongest evidence for the architecture, though, does not depend on Jev at all. TypeSafe’s workflow evals run four business tasks two ways: as a single prompt that asks the model to apply a whole policy, and as a workflow where the policy is broken into narrow typed questions and code combines the answers. TypeSafe reports that “averaged across the four example tasks, every model is more accurate, cheaper and faster in the workflow than it is with the same policy as a prompt.” The model TypeSafe labels haiku 4.5 went from 18.1 percent accuracy as a prompt to 53.6 percent as a workflow. Opus 5 went from 64.8 to 73.1 percent while its cost per case roughly halved. These are vendor-run numbers against vendor-chosen reference labels, and they should be read that way. But they point at a general lesson: separating judgment from control flow helps whichever model does the judging.
One honest gap belongs here. No public census says what fraction of production LLM calls are classifications. The claim that “a large share” are decisions rests on the shape of common architectures, not on measured traffic. A dataset of real API traffic broken down by task type would test it.
System One models vs LLMs
The differences are not only speed and price. They change what a developer has to build around the model.
| Dimension | LLM | System One model (Jev, per TypeSafe) |
|---|---|---|
| Output | A string, parsed afterwards | A typed value from an answer space defined before the call |
| Answer space | Open; schemas can constrain the format | Closed; Choice supports up to 255 options, per the launch post |
| Sampling | Sequential, one token at a time | All answers returned together, “in parallel”, per TypeSafe |
| Uncertainty | Not returned by default; self-reported confidence tends to be overconfident, per TypeSafe | A probability distribution with every answer, and a confidence score on Choice and Score |
| Cost basis | Input and output tokens, output usually dearer | $0.042 per million input tokens; output free |
| Latency | TypeSafe cites 3 to 329 seconds end to end for frontier models | TypeSafe cites 70 to 500 milliseconds |
| Main failure | Fluent, confident, wrong text; malformed output | A valid but wrong label, with a probability that may or may not be honest |
| Flexibility | Writes, codes, explains, reasons over many steps | Cannot generate text; struggles with arithmetic, dates and multi-hop indirection, per its own docs |
| Customization | Prompting, fine-tuning on many platforms | Prompt-side only; the docs say the same weights serve every account |
| Evidence base | Years of public benchmarks and third-party evals | Vendor evals, press anecdotes, no public calibration data |
Two rows deserve a caution. The latency figures vary by who is quoting them. The launch post says 70 to 500 milliseconds, DCVC’s announcement says “less than 100 milliseconds”, and The Register reported that TypeSafe says “as little as 150 ms”, with one demo taking about 620 milliseconds. The cost multiples also depend on the comparison. TypeSafe’s blog says its “193.6x faster, 444.6x cheaper” headline comes from the workflow evals and that it expects “these are on the higher end of real world gains.”
The case for splitting decision from generation
The first argument is cost, and it compounds. An agent that asks a model to choose a tool at every step pays for that choice on every step. The reasoning-model cost problem is that thinking tokens are billed as output, so a decision routed through a reasoning model can cost far more than the answer justifies. On TypeSafe’s aggregate eval chart, Jev’s workflow runs cost $0.0004 and took 0.4 seconds per case. The strongest LLM configuration on the same chart, labelled sol, cost $0.0836 and took 23.3 seconds. A gap that size changes which checks a team can afford to run on every request.
The second argument is latency. A decision that takes tens of seconds cannot sit inside a request path; one that takes a few hundred milliseconds can, which makes a judge on every agent step realistic.
The third argument is the one that matters most: a decision model can say “I am not sure.” The house view, set out in this site’s editorial on hallucination, is that models guess because training and grading reward guessing, and that the remedy is abstention. A probability attached to every answer is an abstention mechanism built into the interface. TypeSafe’s confidence documentation shows the intended use: act automatically on high confidence, confirm on medium, route to a person or a reasoning model on low, with stricter thresholds for destructive actions. The practical guide to Jev walks through setting those thresholds. That is how dependable automation should work, and it is much harder to build on top of a string.
The fourth argument is testability. A typed output can be counted, logged, compared against a label, and regression-tested across model versions. TypeSafe’s FAQ makes the same point, saying “System One tasks are much easier to evaluate.” That claim is right, and it is also the standard the category should be held to.
Calibration has to be shown, not stated
Calibration is the whole product. Speed and price are useful, but the thing that lets code act without a human is a probability that matches reality. If Jev says 0.9, the answer needs to be right about nine times in ten across similar cases, or every threshold built on it is miscalibrated too.
There is good reason to take the goal seriously. The GPT-4 Technical Report found that the pre-trained model was well calibrated, but “after the post-training process, the calibration is reduced”, and its Figure 8 caption says post-training “hurts calibration significantly.” That is the problem RLCD claims to address. There is also good reason for skepticism about any model’s native probabilities. Guo and colleagues showed in 2017 that “modern neural networks, unlike those from a decade ago, are poorly calibrated”, and that a simple post-hoc fix, temperature scaling, often repairs it. Calibration is measurable, fixable and easy to get wrong, which is exactly why it has to be measured.
So far it has not been, in public. TypeSafe’s documentation states that calibration “is measured across groups of predictions; it does not guarantee that an individual answer is correct.” That is accurate and responsible. But the published evidence does not include a reliability diagram, an expected calibration error figure, or any other measurement of whether Jev’s 0.8 means 80 percent. The workflow evals measure accuracy, and accuracy against a proxy: the reference labels are “an average of the responses of GPT-6 Astra and Claude Fable 5.1, both at high thinking.” Agreement with two large LLMs is a reasonable yardstick for a new category. It is not ground truth, and it says nothing about calibration.
TypeSafe’s own docs contain a small example of why this matters. On the ticket “I was charged twice for the same order”, Jev gave 0.72 to “is the customer asking for a refund?” and 0.47 to “is the customer asking for something other than a refund?” The two sum to 1.19. The jaggedness page explains that separate questions are not guaranteed to be consistent and tells developers to “word questions to mean directly what you want.” That is fair engineering advice. It also means a developer cannot assume the numbers behave like probabilities in the textbook sense without testing them on their own data.
Two details raise the stakes. The jev-latest alias “moves when a new release ships, so the answers behind it can change without a change on your side”, according to the models page, which tells teams to pin a version if they have tuned thresholds. And TypeSafe itself argues, in a post on benchmarks, that users should “run your own private evals.” Both are sensible. Together they mean the burden of proving calibration currently sits with every customer, on every version.
A typed answer can still be confidently wrong. The schema guarantees the shape of the answer, not its truth.
“Cannot hallucinate” is the wrong promise
The launch post says Jev “can’t hallucinate.” The reason given is structural: the possible outputs are “defined in advance”, so the model “never makes type errors.” In its hallucination chart, TypeSafe adds a candid note: “Our number is not empirical. Schema matching is guaranteed.”
That guarantee is real and useful. Jev cannot invent a queue that does not exist or return a malformed payload that crashes a pipeline. But it is a guarantee about type, not about truth. The 2025 paper Why Language Models Hallucinate by Kalai, Nachum, Vempala and Zhang argues that the errors people call hallucinations “originate simply as errors in binary classification.” A model that only classifies has not escaped that error. It has changed its form. Instead of a plausible fabricated sentence, the failure is a valid label that happens to be wrong.
Armin Ronacher, CTO of Earendil, put it plainly to TechCrunch: “At the end of the day, it delegates the hallucination problem a little bit to the user.” The jaggedness page lists how the wrong-but-valid answer arises: Jev “answers the question you wrote, not the one you meant”; it “does not count reliably”; it reads dates “as text, not as ordered quantities”; and content written to steer it, such as an injected instruction, “can move the answer.”
There is a second-order risk. A wrong paragraph from a chatbot at least looks like something to check. A clean enum with a high probability attached looks like a fact. Kahneman’s lecture made a version of this point about people: the high error rate on an easy puzzle “illustrates how lightly the output of System 1 is monitored by System 2.” Fast, confident judgment is where unexamined errors live. The category’s honest promise is narrower and better: every answer comes with a number, and the number is worth testing.
The strongest objections
LLMs with structured outputs already do this
Partly true. OpenAI’s Structured Outputs guide says the feature “ensures the model will always generate responses that adhere to your supplied JSON Schema”, including not “hallucinating an invalid enum value.” That removes most of Jev’s type-safety advantage. Research also shows LLM probabilities can carry signal: Kadavath and colleagues found that larger models are “well-calibrated on diverse multiple choice and true/false questions” when the questions are posed in the right format. What structured outputs do not change is the sequential cost and latency of generation, or the effect of post-training on calibration. TypeSafe’s blog notes that the LLMs in its evals used a wrapper that forces decisions with probabilities, and that this “tends to be slower and more expensive.” The objection shows the interface can be copied. It does not show the economics can.
Small fine-tuned classifiers already exist and are cheap
This is the strongest objection. Bucher and Martini found in 2024 that “smaller, fine-tuned LLMs (still) consistently and significantly outperform larger, zero-shot prompted models in text classification.” A fine-tuned BERT-style classifier is fast, cheap, private and can be calibrated with temperature scaling. The counter is that fine-tuning needs labeled data for every task and retraining whenever the labels change, and many decisions in agent systems are defined by a sentence of policy written that morning. A System One model aims to be a zero-shot classifier with the reach of a frontier model. For a stable, high-volume task with labels, a fine-tuned classifier may still be the better choice, and nothing published so far shows otherwise.
A proprietary early-access API is a risky dependency
Also fair. Jev is in early access. The models page warns that rate limits “can change without notice”, and TechCrunch reported that the company “briefly lost the ability to serve users from its API because demand was so high.” The documentation reviewed for this piece describes only a hosted API, and says Jev is not fine-tuned on customer data. The pricing is new, and TypeSafe’s blog concedes, “We can’t prove it isn’t subsidized.” TechCrunch reported that the architecture is undisclosed and that outside observers suspect it is built on an open-weight LLM. Early adopters should put the decision behind an interface they own, so the model underneath can be replaced by a fine-tuned classifier or a constrained LLM if the terms change.
LLM generality wins in the end
General models keep getting cheaper, and one model is simpler to operate than two. On TypeSafe’s evals Jev did not win on accuracy: its 67.8 percent aggregate trailed the configurations labelled sol (74.1 percent) and opus 5 (73.1 percent), and on invoice processing it scored 61.8 percent against sol’s 79.1. The honest reading is that Jev buys cost and speed with some accuracy on harder tasks. That is a routing problem, not a refutation. A confident Jev answer can act immediately; an uncertain one can escalate to a reasoning model. Generality and specialization are not rivals in that design. They are layers, and test-time compute should be spent only where the fast layer admits it is unsure.
What would change the conclusion
Several findings would change this view, in either direction.
- Independent calibration data. Reliability diagrams and calibration error figures for Jev on real workloads, published by someone other than TypeSafe, per model version. If they show Jev’s confidence tracks accuracy, the category has its proof. If they show it does not, the main reason to prefer it over a constrained LLM is gone.
- Head-to-head against fine-tuned classifiers. A fair comparison on labeled tasks, including cost of labeling and retraining, would show where the zero-shot convenience is worth paying for.
- Robustness under attack. Measured results on prompt injection in the state, which the jaggedness page names as a known weakness.
- Durable economics and access. Stable rate limits, a pricing history, and some path for buyers who cannot send data to a new vendor’s hosted API.
- A census of real traffic. If measured production workloads turn out to be mostly generation after all, the case for a separate decision layer shrinks to a niche.
The Register quoted YouTuber Mo Bitar on Jev: “I know it’s fast. I know it’s cheap, but is it good?” That is the right question, and it has a specific answer waiting to be measured. “Good”, for a decision model, means its probabilities are honest. The split between deciding and talking is overdue. The proof that System One models decide honestly is still to come.
Sources
- Introducing System One Models & Jev (TypeSafe AI, 15 September 2026)
- System One (TypeSafe documentation)
- AI primer (TypeSafe documentation)
- Confidence (TypeSafe documentation)
- Models (TypeSafe documentation)
- Jev 1.13 jaggedness (TypeSafe documentation, reviewed 17 September 2026)
- Workflow evals (TypeSafe AI)
- Lies, Damned Lies, and Benchmarks (TypeSafe AI, 11 September 2026)
- TypeSafe emerges from stealth with a new way of doing AI (DCVC, 15 September 2026)
- A new kind of AI model from a ChatGPT inventor is thrilling developers (Fernholz, TechCrunch, 18 September 2026)
- Shut up and calculate: Jev's new AI primitives for coders (Jackson, The Register, 23 September 2026)
- Building a harness with Jev (Runkle and Lovell, LangChain, 17 September 2026)
- Maps of Bounded Rationality: A Perspective on Intuitive Judgment and Choice (Kahneman, Nobel Prize Lecture, 8 December 2002)
- GPT-4 Technical Report (OpenAI, 2023)
- On Calibration of Modern Neural Networks (Guo, Pleiss, Sun and Weinberger, 2017)
- Language Models (Mostly) Know What They Know (Kadavath et al., 2022)
- Why Language Models Hallucinate (Kalai, Nachum, Vempala and Zhang, 2025)
- Fine-Tuned 'Small' LLMs (Still) Significantly Outperform Zero-Shot Generative AI Models in Text Classification (Bucher and Martini, 2024)
- RouteLLM: Learning to Route LLMs with Preference Data (Ong et al., 2024)
- Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena (Zheng et al., 2023)
- Structured Outputs guide (OpenAI API documentation)
Questions readers ask
What are System One models?
TypeSafe AI defines System One models as "a class of AI models built to make fast, structured decisions that software can use directly." A System One model reads a state and returns typed answers with probabilities rather than generated text. Jev, released in early access on 15 September 2026, is the first one.
What is the difference between a System One model and an LLM?
An LLM generates a string one token at a time and is billed for input and output tokens. A System One model returns a value from an answer space the developer defines in advance, with a probability distribution, and TypeSafe bills Jev only for input tokens at $0.042 per million. It cannot write replies, code or explanations.
Can a System One model hallucinate?
It cannot return an answer outside the schema, which is what TypeSafe's "can't hallucinate" claim rests on. It can still pick the wrong option with high confidence. TypeSafe's own documentation says calibration "does not guarantee that an individual answer is correct."
Is there independent evidence that Jev is calibrated?
Not in public as of 29 September 2026. TypeSafe's published workflow evals measure agreement with reference labels from two large LLMs, and early users have reported speed and accuracy anecdotes to the press. No independent reliability diagram or calibration error figure for Jev has been published.
