Independent reporting on artificial intelligence.


The Frontier Wire

Research

System One models, explained: typed answers and calibration

TypeSafe's Jev returns typed answers with probabilities instead of text. The idea borrows Kahneman's System 1, a calibration objective called RLCD, and a strong claim about hallucination. The concepts, the evidence behind them, and where each one stops.

TL;DR

  • A System One model, in TypeSafe’s definition, “evaluates a state and returns typed answers and probabilities.” It does not generate text. The developer fixes the answer space in advance as a Choice, a Score or a yes/no “Noul”.
  • The name borrows Kahneman’s System 1, the fast and automatic mode of thought. The analogy fits the intended workload of quick, narrow judgments. It breaks on error: Kahneman’s System 1 is the source of predictable mistakes, and TypeSafe claims its models are calibrated.
  • Calibration means that answers given with 80 percent confidence are right about 80 percent of the time. It is measured over many predictions with reliability diagrams and expected calibration error. TypeSafe describes its training objective, RLCD, as optimizing for it, but has not published a reliability diagram or ECE figure for Jev.
  • “Cannot hallucinate” holds only in a narrow sense: Jev cannot return a value outside the options you supplied. It can still choose the wrong option, which TypeSafe’s own docs acknowledge.

On 15 September 2026, TypeSafe AI released a model called Jev and described it as the first of a new class it calls System One models. The label packs several ideas into two words: a reference to a famous theory of human thinking, a different output contract from the one large language models use, and a training objective built around calibrated probability. This explainer takes those ideas one at a time, sets out what each means in general terms, and marks where TypeSafe’s descriptions are company claims rather than independently established facts. The desk has not run Jev. The news of the launch is covered in our report on the release, and practical usage in our guide to using Jev.

The short definition

TypeSafe’s documentation gives the definition in one sentence: “A System One model evaluates a state and returns typed answers and probabilities.” The System One concept page adds that “like an LLM, a System One model understands natural-language input. It returns typed decisions and probabilities rather than generated text.”

Two terms carry the definition. The state is the input: a string, a JSON object or an array of text that describes the situation to be judged, such as a support ticket plus the customer’s order history. A typed answer is an output whose possible values are fixed before the model runs. TypeSafe offers three question types, which it calls primitives:

  • Choice picks one option from a list the developer supplies and returns “the selected option, a probability for each option, and confidence.” A single Choice accepts up to 255 options, according to the Choice reference.
  • Score rates the state against ordered levels the developer describes, such as calm, frustrated and very frustrated, and returns a position along them, a probability for each level and confidence.
  • Noul answers a yes-or-no question with “the probability that the answer is yes.” It has no separate confidence field, because a two-outcome distribution is fully described by one number.

Several questions can be sent against one state in a single request. The introduction page says every question “is evaluated in parallel and in isolation against the same state in one go.”

What a System One model does not do is equally specific. The docs say System One models do “not write replies, produce code, or generate explanations of their reasoning.” A separate page for developers arriving from coding tools states that Jev is “not a drop-in replacement for the LLM behind Claude Code, Cursor” and similar agents. The Jev 1.13 limitations page lists generation as a failure mode and advises: “If you really need to generate text… there are other models for that.”

System 1 and System 2 in Kahneman’s account

The name comes from Daniel Kahneman’s dual-process account of thinking, popularized in Thinking, Fast and Slow. TypeSafe’s docs cite the book directly: “System 1 thinking is fast and intuitive. System 2 is slower and more deliberate. Here, the emphasis is on fast, focused judgments.”

Kahneman set out the framework in more technical terms in his 2002 Nobel Prize lecture, Maps of Bounded Rationality. There he credits the labels to the psychologists Keith Stanovich and Richard West, and describes the two systems this way: “The operations of System 1 are fast, automatic, effortless, associative, and difficult to control or modify. The operations of System 2 are slower, serial, effortful, and deliberately controlled; they are also relatively flexible and potentially rule-governed.”

The lecture also gives System 2 a supervisory job. “One of the functions of System 2 is to monitor the quality of both mental operations and overt behavior,” Kahneman writes, and he adds that “the monitoring is normally quite lax, and allows many intuitive judgments to be expressed, including some that are erroneous.” His example is a puzzle: a bat and a ball cost $1.10 together, and the bat costs $1 more than the ball. The fast answer, 10 cents, comes to mind first and is wrong.

Where the analogy holds

Mapped onto AI systems, the System 2 side is easy to identify. Reasoning models spend extra computation at inference time, generating a long internal chain of thought before they answer. OpenAI’s reasoning guide says these models “use internal reasoning tokens before producing a response,” which “helps the model plan, use tools effectively, inspect alternatives, recover from ambiguity, and solve harder multi-step tasks.” Those tokens are billed as output. Our explainer on test-time compute covers the research behind that approach, and our report on reasoning costs covers the bill. In Kahneman’s terms this is slow, serial and effortful: more thinking per request, at more cost per request.

A System One model sits at the other end. Three parts of Kahneman’s description line up with TypeSafe’s design.

Speed. TypeSafe’s launch post claims end-to-end response times of “70ms-500ms”, against “3 to 329 seconds” it cites for frontier models. Its build guide says “most queries complete in about 100 ms.” These are the company’s figures.

Parallel rather than serial. Kahneman calls System 2 serial. An LLM also works serially, generating one token at a time, each conditioned on the last. TypeSafe says its model “generates all outputs in a single query,” with every question evaluated in parallel.

Narrow, immediate judgments. TypeSafe tells developers to ask for “a judgment a knowledgeable person makes in a second given the right context.” Its example of a good question is “Does this message convey urgency?” Its example of a bad one is “Analyze this message and determine the best course of action,” which, the docs say, “needs slow reasoning, and it is a signal to break the task into small questions and compose the answers in code.”

That last point is where the analogy is most useful. In the design TypeSafe describes, the developer’s code plays the monitoring role that Kahneman gives System 2. The model supplies fast judgments. The code decides which ones to trust, combines them, and sends uncertain cases elsewhere. The concept page says so directly: answers “include confidence, so you can decide when to act and when to escalate to a person or a reasoning model.”

Where the analogy breaks

The analogy fails at the point Kahneman cared about most. His research program was about the systematic errors of intuition: anchoring, framing, substituting an easier question for a harder one. In his account, System 1 is fast because it skips checks, and the result is predictable mistakes.

TypeSafe acknowledges the tension. The launch post’s FAQ notes that “‘System 1 thinking’ has also implied error-prone” and says the company believes “System One Models can be made more reliable than its alternatives,” with the reasons deferred to future posts. The claim that separates a System One model from human System 1 is calibration: a person’s snap judgment carries no reliable measure of its own uncertainty, while TypeSafe says its model’s probabilities do. Whether that holds is an empirical question, and it is taken up below.

There are two further mismatches. First, human System 1 is associative and hard to direct. A System One model is directed completely: it answers exactly the question written, over exactly the options supplied. TypeSafe’s limitations page warns that jev-1.13 “answers the question you wrote, not the one you meant,” which is closer to a literal-minded function than to intuition. Second, Kahneman’s systems are two modes of one mind. In the TypeSafe design they are separate components: one model for fast judgments and, where needed, a different model or a person for deliberation.

Typed outputs versus generated text

A language model produces text. When software needs a decision from it, the text has to be turned back into a value. TypeSafe’s introduction describes the problem as “coercing a text-generation system into outputting structured decisions, then parsing the results back into something your code can depend on.”

LLM providers have already narrowed that gap, in stages.

JSON mode, structured outputs and constrained decoding

JSON mode guarantees that the model’s output parses as JSON, but not that it has the fields the application expects. Structured Outputs, in OpenAI’s API, goes further. The OpenAI guide says the feature “ensures the model will always generate responses that adhere to your supplied JSON Schema, so you don’t need to worry about the model omitting a required key, or hallucinating an invalid enum value.” Its comparison table is explicit: both modes output valid JSON, and only Structured Outputs adheres to the schema.

The general technique underneath is constrained decoding: at each step of generation, tokens that would break the required format are masked out. Willard and Louf’s 2023 paper, Efficient Guided Generation for Large Language Models, reformulates generation “in terms of transitions between the states of a finite-state machine,” so that output can be restricted to a regular expression or a context-free grammar. Their open-source library, Outlines, implements it.

Even with a schema enforced, the guarantee has edges. OpenAI’s guide warns that “the model might not generate a valid response that matches the provided JSON schema” in the case of a safety refusal or when a max-token limit cuts the response short. Refusals come back in a separate refusal field.

Constrained generation and System One scoring guarantee the same kind of thing, the shape of the answer, and neither guarantees that the answer is correct. They differ in three ways.

  1. What comes back. Structured Outputs returns one chosen value. A System One model returns a probability for every allowed value. An application can get probabilities from some LLM APIs through token log-probabilities, but that is extra work and depends on how the answer was tokenized.
  2. How the work scales. An LLM generates its structured answer token by token. TypeSafe says Jev scores all questions in parallel and that “adding questions barely changes the response time.”
  3. What it costs. LLM APIs bill generated output tokens, usually at a higher rate than input. TypeSafe’s models page lists Jev 1.13 at $0.042 per million input tokens, with “output tokens are free.”

TypeSafe itself uses an LLM with structured output as its point of comparison. Its published evaluations run competing models through what it calls a “System One LLM wrapper,” which “constrains LLMs to output structured decisions compatible with our API.”

Calibration, defined

The word that carries most of TypeSafe’s argument is “calibrated.” It has a precise meaning.

The standard reference in deep learning is Guo, Pleiss, Sun and Weinberger’s On Calibration of Modern Neural Networks, presented at ICML in 2017. It defines calibration as the property that a model’s confidence “represents a true probability.” Their example: “given 100 predictions, each with confidence of 0.8, we expect that 80 should be correctly classified.” Formally, perfect calibration means that for every confidence level p, the chance of being correct given confidence p equals p. The authors add that “in all practical settings, achieving perfect calibration is impossible,” so it has to be estimated.

TypeSafe’s AI primer uses the same definition: across many predictions from a well-calibrated model, outcomes assigned 0.2 “should occur about 20% of the time,” and outcomes assigned 0.8 about 80 percent of the time.

Reliability diagrams and expected calibration error

Guo and colleagues describe two tools for measuring calibration.

A reliability diagram sorts predictions into bins by confidence, for example 0 to 0.1, 0.1 to 0.2 and so on, and plots each bin’s actual accuracy against its average confidence. “If the model is perfectly calibrated,” the paper says, “the diagram should plot the identity function. Any deviation from a perfect diagonal represents miscalibration.” Bars below the diagonal mean overconfidence. Bars above it mean underconfidence. The paper also notes a limit: reliability diagrams “do not display the proportion of samples in a given bin.” A model can look well calibrated in a bin that holds only a handful of predictions.

Expected calibration error (ECE) reduces the diagram to one number. It takes the absolute gap between accuracy and confidence in each bin and averages those gaps, weighting each bin by the share of predictions in it. Zero is perfect. The paper also defines maximum calibration error, the largest gap in any bin, for “high-risk applications where reliable confidence measures are absolutely necessary.”

Two further points from the calibration literature matter for anyone reading a vendor’s claim.

First, calibration is not accuracy. The scikit-learn documentation notes that proper scoring rules such as the Brier score mix together calibration and “discriminative power (resolution),” and that a lower score “does not necessarily mean a better calibrated model.” A model that always predicts the base rate can be perfectly calibrated and useless. The useful model is both calibrated and sharp, meaning it gives confident answers where it can.

Second, calibration can be adjusted after training without changing the answers. Guo and colleagues found that temperature scaling, which divides the model’s raw scores by a single learned constant before converting them to probabilities, was “surprisingly effective,” and because it does not change which option scores highest, “temperature scaling does not affect the model’s accuracy.”

Why LLM probabilities drifted

Guo’s paper opened with a finding that surprised its authors: “modern neural networks, unlike those from a decade ago, are poorly calibrated.” A 110-layer ResNet was more accurate than a 5-layer LeNet from 1998 on the CIFAR-100 image dataset, but its “average confidence … is substantially higher than its accuracy.”

Language models showed a related pattern. Kadavath and colleagues at Anthropic reported in 2022, in Language Models (Mostly) Know What They Know, that “larger models are well-calibrated on diverse multiple choice and true/false questions when they are provided in the right format.” Then the GPT-4 Technical Report measured what post-training did to that property. The pre-trained GPT-4 was “highly calibrated,” the report says, but “after the post-training process, the calibration is reduced.” Its Figure 8, on a subset of the MMLU benchmark, shows an ECE of 0.007 for the pre-trained model and 0.074 for the post-trained one, with the caption: “The post-training hurts calibration significantly.”

This is the problem TypeSafe says it set out to fix. Its primer argues that reinforcement learning from human feedback (RLHF) “teaches a model to say things that people prefer,” which “can also reward sycophancy and confident-sounding hallucinations.” The GPT-4 result supports part of that argument: at least one widely used post-training process made a well-calibrated model less so. It does not show that a different objective fixes the problem. That requires evidence about the new model.

What TypeSafe discloses about RLCD

TypeSafe calls its training method reinforcement learning for calibrated decisions, or RLCD. (TechCrunch’s report renders it as “reinforcement learning from calibrated decisions”; TypeSafe’s own pages use “for.”) The primer places RLCD as a third post-training path after two established ones. RLHF “turned pretrained models into chatbots.” Reinforcement learning with verifiable rewards (RLVR) “created reasoning models that are strong at tasks such as mathematics, but slower and more expensive.” RLCD “trains TypeSafe to return decisions and calibrated probabilities instead of generated text.”

The primer states the output contract RLCD optimizes for in three lines: “The model does not generate text. It returns decisions and probabilities. Higher probability should correspond to a greater chance that the answer is correct.” The concept page adds that probabilities “are optimized against outcomes to reflect uncertainty,” and then adds a qualification: “Calibration is measured across groups of predictions; it does not guarantee that an individual answer is correct.”

Beyond that, the public record is thin:

  • Architecture. The launch post says TypeSafe built “a new model architecture, parallel sampler for maximum efficiency,” and training method. It does not describe the architecture. TechCrunch calls Jev “transformer-based” and reports that Almeida “is tight-lipped about the model’s architecture, which outside observers suspect is built on top of an open-weight LLM.”
  • Model size. Not disclosed. Asked in the launch FAQ whether Jev is “just a smaller LLM,” TypeSafe answers: “Jev is neither small nor an LLM.”
  • Training data. The launch FAQ says “we make all the data ourselves.” Almeida told TechCrunch that Jev “is trained exclusively on synthetic data.” The models page says Jev “is not trained on customer requests or responses” and “is not fine-tuned or LoRA-adapted with customer data”; the same weights serve every account.
  • Reward. No public description of the reward function, the outcomes probabilities are scored against, or the loss used.
  • Calibration evidence. None of the TypeSafe pages consulted for this article publishes a reliability diagram, an ECE figure or a comparable calibration measurement for Jev.

The last point matters most. A calibration claim is testable, and the test is standard. TypeSafe has said it will not publish results on public benchmarks. Its post on benchmarks argues that public evaluations get gamed and commits instead to “dated snapshots” and internal evaluations published “alongside the caveats.” That is a defensible position on leaderboards. It leaves each buyer to measure calibration on their own data.

The founder’s background explains the framing. Diogo Almeida is the fourth of twenty authors on OpenAI’s InstructGPT paper, the 2022 work that applied RLHF to instruction following. TypeSafe’s docs say he “co-invented” RLHF. TechCrunch’s report says he helped “invent” it.

What “cannot hallucinate” can and cannot mean

TypeSafe’s launch post says Jev “gives up string generation, it’s optimized for structured outputs and can’t hallucinate.” TechCrunch’s report explains the reasoning: “because users define the outputs in advance, it cannot hallucinate.”

Whether that is true depends on the definition of hallucination. Our editorial on the word argues that it covers several different failures, and the distinction helps here.

What the claim does cover. Jev cannot return a value that is not one of the options the developer wrote. The primitives page states that answers are “constrained to the options you supplied,” with “never a value outside them.” It cannot invent a nonexistent category, return a misspelled enum or produce unparseable output. TypeSafe calls this the no-type-error guarantee. In the launch post it says a type error “would be an easy thing to falsify with just a single counter-example, but it is mathematically impossible.” For the hallucination-rate chart in that post, TypeSafe notes that “our number is not empirical. Schema matching is guaranteed, thus we can confidently add 0% into the plots.” That is a property of the output format, not a measured error rate.

What the claim does not cover. A typed answer can be wrong. If a refund request is routed to “technical” instead of “billing,” the answer is well formed and incorrect. TypeSafe’s own documentation says so repeatedly. The concept page says calibration “does not guarantee that an individual answer is correct.” The limitations page for Jev 1.13 lists nine failure modes, including literal reading, unreliable counting, date comparison, multi-hop indirection, distraction by irrelevant state and adversarial content that “can move the answer.”

The same page shows a subtler case. Asked “Is the customer asking for a refund?” and its negation as two separate Nouls about the same ticket, Jev returned 0.72 and 0.47. Those add to 1.19, where a coherent pair would add to 1. TypeSafe’s guidance is not to “hold the model to arithmetic identities between separate questions.” Each answer may be individually reasonable while the pair is inconsistent.

Armin Ronacher, CTO of Earendil, described the trade in TechCrunch’s report: “At the end of the day, it delegates the hallucination problem a little bit to the user. The user has to say, okay, if this only comes back with 50% probability, maybe this is a coin toss, and I disregard it. But if it’s 95%, sure, then I can do something with it.”

That is the accurate version of the claim. A System One model turns invented values into wrong choices, and puts a probability next to each choice. Whether the probability can be trusted is the calibration question again.

A typed answer removes the invented value. It does not remove the wrong one.

Confidence is not the same as correctness

One more distinction matters in practice. TypeSafe returns both probabilities and a confidence field on Choice and Score answers. The confidence page says confidence “is a statistic computed from the probability distribution”: 1.0 when all probability is on one option, lower as the distribution flattens. Its interactive example uses “(3 × largest probability − 1) / 2” to approximate confidence for three options. Confidence is therefore a measure of how peaked the distribution is, not a separate estimate of the chance of being right. Its usefulness depends on the probabilities underneath being calibrated. TypeSafe’s own cookbook on uncertainty bands says the thresholds it illustrates are “neither a calibrated guarantee nor an optimized threshold” and tells developers to set production boundaries “from labeled examples.”

Classic classifiers and fine-tuned small models

A system that takes input and returns one of a fixed set of labels with a probability is not new. It is the definition of a probabilistic classifier, and The Register’s report on Jev described it as “a classifier with brains.”

A classic classifier, such as logistic regression or a gradient-boosted tree, is trained on labeled examples of the specific task. It needs that data, and it knows only that task. In exchange, it is small, fast and cheap to run on a team’s own hardware, and its calibration can be measured and corrected with standard tools. The scikit-learn documentation notes that logistic regression “is more likely to return well calibrated predictions by itself,” while other models “return biased probabilities; with different biases per model.” It provides a CalibratedClassifierCV wrapper that fits a correction on held-out data.

A fine-tuned small language model sits in between: a pretrained model adapted to a task with a smaller labeled set. It keeps some general language understanding and still needs task data, training runs and hosting.

The System One pitch is to keep the output contract of a classifier while dropping the per-task training. The label set is written into the request, not learned from examples. The models page states that customization happens “through the request rather than through per-account weights”: proprietary data goes in the state, and domain rules go in the instructions and criteria. TypeSafe’s AutoResearch cookbook goes the other way and uses Jev’s probabilities as features for a downstream CatBoost regressor, a classic model trained on top.

The trade is familiar from LLMs generally. A zero-shot model avoids the cost of collecting labels but gives up the guarantee that it has learned the team’s exact decision boundary. The labeled data is still needed, to evaluate the model, even if it is not used to train it.

Four approaches side by side

The table compares the four approaches discussed above on the criteria a team would weigh for a classification-style decision. Entries for System One models reflect TypeSafe’s published documentation and claims, not independent measurement.

CriterionClassic classifierLLM with structured outputReasoning modelSystem One model (Jev)
OutputA label and a probability per classGenerated text constrained to a JSON schema; one value, no probability by defaultGenerated text after hidden reasoning tokens; can also be schema-constrainedA typed Choice, Score or Noul with a probability for every allowed value
LatencyRuns locally; no network round trip neededOne generation pass, token by tokenLongest; scales with reasoning lengthTypeSafe claims 70 to 500 ms end to end
Cost basisOwn compute for training and servingInput plus generated output tokensInput plus output, with reasoning tokens billed as output$0.042 per million input tokens; output free
CalibrationMeasurable and correctable with standard tools; varies by model typeNot exposed by default; post-training reduced it in GPT-4’s caseNot exposed by defaultClaimed as the training objective (RLCD); no published ECE or reliability diagram
Training data needLabeled examples for each taskNone to start; labeled data for evaluationNone to start; labeled data for evaluationNone to start; labeled data for evaluation and threshold setting
FlexibilityOne task per modelAny task expressible as textAny task, including multi-step reasoningFixed-answer judgments only; up to 255 options per Choice; no text generation

TypeSafe’s own evaluation gives one view of where the model lands against LLMs. Its workflow evals site averages four business workflows, scored against “consensus labels” produced by GPT-6 Astra and Claude Fable 5.1 at high thinking settings, not against human ground truth. On that average Jev scores 67.8 percent at $0.0004 and 0.4 seconds per case. GPT-6 Sol scores 74.1 percent at $0.0836 and 23.3 seconds, and Claude Opus 5 scores 73.1 percent at $0.1761 and 37.8 seconds. On invoice processing, Jev scores 61.8 percent against 79.1 percent for Sol. The picture TypeSafe’s own numbers draw is lower agreement with the reference models than the strongest reasoning LLMs, at a small fraction of the cost and time. The launch post also discloses that the workflows “were made by individuals on our model capabilities team, so some bias could exist.”

Claims and reported evidence

The speed and cost figures around System One models vary by source, so they are worth separating.

SourceClaim
TypeSafe launch post70 to 500 ms end to end; “40x-200x faster” on System One-shaped queries; home-page figures of “193.6x faster, 444.6x cheaper,” which the post says “are on the higher end of real world gains”
TypeSafe build guide“Most queries complete in about 100 ms”
DCVC, lead investor in the $40 million seed round“Less than 100 milliseconds of latency”; “often 100 times faster and less expensive”
LangChain blog“The company reports up to 200x faster inference and 400x lower cost” on classification tasks
The Register“TypeSafe says Jev can return an answer in as little as 150 ms”
TechCrunch, developer reportsA Vercel engineer reported results “five to 18 times more quickly” than an OpenAI model on a safety classifier; a Bryo AI tester found Gemini “slightly more accurate, but 10 to 20 times more expensive” on email classification

The developer reports in TechCrunch are the closest thing to independent testing in the public record, and they are anecdotes, not published methodology. The Register quotes one engineer who probed Jev with 10,000 API calls and another commentator who asked the open question directly: “I know it’s fast. I know it’s cheap, but is it good?”

Testing the calibration claim yourself

Because TypeSafe publishes the full probability distribution and no calibration figures, the check falls to the user. The method is the one Guo and colleagues describe, and it needs no special access.

  1. Collect labeled cases. Take a few hundred real inputs for the decision in question and label the correct answer. The cases should reflect the production mix, including the hard ones.
  2. Run the questions. Record the chosen option and its probability for each case. For a Noul, record the probability of yes.
  3. Bin and plot. Group predictions into ten confidence bins. For each bin, compare average confidence with the share actually correct. That is the reliability diagram.
  4. Compute ECE and check the counts. Weight each bin’s gap by its share of predictions. Look at how many predictions fall in each bin; a well-calibrated bin with five cases in it proves little.
  5. Compare against the alternative. Run the same set through the LLM or classifier currently in use, with probabilities extracted the same way, and compare accuracy, ECE, latency and cost together.
  6. Pin the version. The models page warns that the jev-latest alias “can change without a change on your side” and advises pinning a versioned ID such as jev-1.13.0 if thresholds were tuned against it.

Our guide to comparing LLM evaluation frameworks covers tooling for running this kind of test at scale. The underlying point is simple. A calibration claim is only as good as the reliability diagram behind it, and for Jev that diagram has to be drawn by the buyer.

Sources

  1. Introducing System One Models & Jev (TypeSafe, 15 September 2026)
  2. TypeSafe docs: System One
  3. TypeSafe docs: AI primer (RLCD)
  4. TypeSafe docs: Confidence
  5. TypeSafe docs: Primitives (Questions)
  6. TypeSafe docs: Models
  7. TypeSafe docs: Jev 1.13 jaggedness
  8. TypeSafe docs: How to build with TypeSafe
  9. TypeSafe docs: Self-consistency, nouls
  10. TypeSafe workflow evals
  11. Lies, Damned Lies, and Benchmarks (TypeSafe, 11 September 2026)
  12. TypeSafe emerges from stealth with a new way of doing AI (DCVC, 15 September 2026)
  13. Maps of Bounded Rationality: A Perspective on Intuitive Judgment and Choice (Kahneman, Nobel Prize Lecture, 2002)
  14. On Calibration of Modern Neural Networks (Guo, Pleiss, Sun and Weinberger, ICML 2017)
  15. GPT-4 Technical Report (OpenAI, 2023)
  16. Language Models (Mostly) Know What They Know (Kadavath et al., 2022)
  17. Training language models to follow instructions with human feedback (Ouyang et al., 2022)
  18. Efficient Guided Generation for Large Language Models (Willard and Louf, 2023)
  19. OpenAI API: Structured model outputs
  20. OpenAI API: Reasoning models
  21. scikit-learn: Probability calibration
  22. A new kind of AI model from a ChatGPT inventor is thrilling developers (TechCrunch, 18 September 2026)
  23. Shut up and calculate: Jev's new AI primitives for coders (The Register, 23 September 2026)
  24. Building a Harness with Jev (LangChain, 17 September 2026)

Questions readers ask

What is a System One model?

It is TypeSafe's name for a model that takes a state (text or JSON) and a set of questions whose possible answers are defined in advance, and returns a typed answer with a probability distribution instead of generated text. Jev, released in early access on 15 September 2026, is the first one. The name refers to Daniel Kahneman's System 1, the fast, automatic mode of thought.

How is a System One model different from an LLM with structured outputs?

An LLM with structured outputs still generates text token by token, and the API constrains that text to match a JSON schema. The schema is guaranteed but the values are not, and no probability comes back by default. A System One model scores a fixed answer space directly and returns a probability for every option. Both guarantee the shape of the answer; neither guarantees that the answer is right.

What does calibration mean for a model?

A model is calibrated if, across many predictions made with confidence p, about p of them turn out correct. Guo et al. (2017) define it formally and measure it with reliability diagrams and expected calibration error (ECE), the weighted average gap between confidence and accuracy across probability bins. Calibration is a property of groups of predictions, not a guarantee about any single answer.

Can Jev really not hallucinate?

TypeSafe says Jev "can't hallucinate", and the defensible reading is narrow. Because every answer must be one of the options the developer supplied, Jev cannot invent a value outside that set, and it cannot return malformed output. It can still pick the wrong option, and TypeSafe's own documentation says calibration "does not guarantee that an individual answer is correct".

More research