How to use Jev: where a System One model fits next to an LLM
TypeSafe's Jev returns typed answers with probabilities instead of text. Its documentation, the first developer reports and early independent papers show where it fits in a pipeline, what the published price works out to, and why its probabilities need checking on your own data.
TL;DR
- Jev takes a
state(text or JSON) and a map of typed questions, and returns one answer per question: a Choice from your options, a Score on your rubric, or a Noul, the probability that a statement is true. Choice and Score answers also carry aconfidencevalue. It does not generate text.- Early users describe it as a fast decision step inside software: routing requests to the right model or handler, checking agent tool calls, screening LLM inputs and outputs, judging generated text, and classifying records at volume.
- The pattern TypeSafe and LangChain both document is “cheap by default, frontier on exception”: Jev decides, and anything below a confidence threshold goes to an LLM, a reasoning model or a person.
- The published price is $0.042 per million input tokens with free output. That is well below frontier LLM rates, but the gap to the smallest LLMs is a few times, not hundreds.
- TypeSafe says the probabilities are calibrated. Independent write-ups and early preprints say calibration varies by task and question type. Check it on your own labeled data before you set thresholds.
TypeSafe AI released Jev in early access on 15 September 2026 and called it the first “System One model”, a class of model that makes structured decisions instead of writing text. Two weeks later there is enough public material to say how developers are putting it to work: TypeSafe’s API reference and cookbooks, two LangChain engineering posts, reporting in TechCrunch and The Register, and a handful of independent tests and preprints. This guide draws on those sources to explain how the API works, where Jev sits next to an LLM, what the price works out to, and how to test its probabilities before relying on them. The Frontier Wire did not run Jev for this piece. Every result below is attributed to whoever produced it.
For the launch itself, see TypeSafe launches Jev. For the concept, see System One models, explained.
What a System One model does
TypeSafe’s documentation defines a System One model as one “built to make fast, structured decisions that software can use directly”. It “evaluates a state and returns typed answers and probabilities.” The name comes from Daniel Kahneman’s Thinking, Fast and Slow, where System 1 is fast, intuitive thinking and System 2 is slow and deliberate. Jev is named after William Stanley Jevons, the economist associated with the idea that cheaper resources get used more.
Like an LLM, Jev reads natural language. Unlike an LLM, it does not “write replies, produce code, or generate explanations of their reasoning”, in the documentation’s words. The company’s own page for developers who arrive from coding tools is blunt: Jev “is not a drop-in replacement for the LLM behind Claude Code, Cursor” and similar tools. It is a component you call from code.
The design advice that runs through the docs is to keep each question small. TypeSafe calls this “atomic questions, composed in code”. Instead of asking one question that weighs several factors, ask one question per factor and combine the answers in your own logic. The Register’s Joab Jackson summarized the result as “a classifier with brains”, and noted the cost of the approach: “using Jev requires some old-school manual configuration ahead of time”.
The request: state and questions
Every call goes to one endpoint, POST https://api.typesafe.ai/v1/systemone, with a bearer token. The body has three required fields.
stateis the content to evaluate. It can be a string, a JSON object or an array. TypeSafe’s State page recommends an object for most requests “so each part of the state has a descriptive name”, for example a ticket, the order it refers to and the refund policy, together in one state.modelselects the model. The docs usejev-latest, which currently points tojev-1.13.0.questionsis a map of named questions. You choose each key, and answers come back under the same keys. The API reference notes that the key “is not sent to the underlying model”.
The request below is copied from TypeSafe’s quick start. It asks three questions of one support message.
{
"state": "Hi, I've been trying to connect my Stripe account for 3 days and the integration keeps failing. I'm losing sales. Please help ASAP.",
"model": "jev-latest",
"questions": {
"department": {
"type": "choice",
"instructions": "Which team should handle this",
"criteria": {
"billing": "Payment or subscription issues",
"technical": "Bugs or integration problems",
"sales": "Pricing or account questions"
}
},
"frustration": {
"type": "score",
"instructions": "How frustrated the customer appears",
"criteria": [
"Calm, just stating facts",
"Frustrated but civil",
"Very angry, strong language"
]
},
"is_urgent": {
"type": "noul",
"instructions": "The message conveys urgency or time-sensitivity"
}
}
}
Declaring the output type
The type field on each question is how you declare what shape of answer you want. There are three.
| Type | Use it when | You define | Limits in the API reference |
|---|---|---|---|
| Choice | The answer is one of a known, unordered set: a department, a document type, a tool | A map of option names to descriptions | Up to 255 options |
| Score | The answer sits on a spectrum you can describe: severity, frustration, skill | An ordered array of level descriptions | 2 to 10 levels |
| Noul | A clean yes or no, where the probability itself is useful | The statement, plus optional meanings for true and false | Returns one value from 0 to 1 |
TypeSafe’s guidance on choosing is practical: “prefer the one whose answer your code can act on directly.” A Choice maps onto code paths, a Score onto a threshold, a Noul onto an if. The docs also warn against a common confusion. A Noul of 0.5 “means the model gives yes and no equal probability. It does not mean the candidate has a medium skill level.” If you want a level, use a Score.
For Choice questions, the docs recommend adding an other or none of the above option whenever the list might not cover every input. Otherwise the model must put its probability somewhere among options that do not fit.
All questions in a request are evaluated in parallel and independently against the same state. TypeSafe says adding questions “barely changes the response time” and costs only the tokens for the extra questions. Its parallel questions cookbook ran 13 questions over a 53,777-character GDPR article and reported that one batched call was 12.2 times cheaper and 10 times faster than 13 single-question calls, with no change in the answers. The reason is simple: every separate call pays for the document again.
The response: answers, probabilities and confidence
The quick start’s sample response to the request above looks like this (again reproduced from TypeSafe’s documentation):
{
"model": "jev-1.13.0",
"answers": {
"department": {
"type": "choice",
"choice": "technical",
"confidence": 0.78,
"probabilities": { "technical": 0.85, "sales": 0.0, "billing": 0.15 }
},
"frustration": {
"type": "score",
"score": 1.0,
"confidence": 1.0,
"legend": { "0": "Calm, just stating facts", "1": "Frustrated but civil", "2": "Very angry, strong language" },
"probabilities": { "0": 0.0, "1": 1.0, "2": 0.0 }
},
"is_urgent": { "type": "noul", "noul": 1.0 }
},
"usage": { "input_tokens": 392, "output_tokens": 65 }
}
Three things are worth reading closely.
Probabilities. Choice and Score answers return the full distribution over your options or levels, summing to 1. A Choice’s choice is simply the highest-probability option. A Score’s score is “the probability-weighted answer across the levels” and “can land between levels”, so a 1.4 is possible on a 0 to 2 scale.
Confidence. Choice and Score answers also carry confidence, a number from 0 to 1 that TypeSafe computes from how concentrated the distribution is. All the probability on one option gives 1.0. An even spread gives a low value. Noul answers do not carry a confidence field; the Noul value is itself the probability. The docs say you are “never locked into our definition” and can compute your own measure from the returned probabilities.
The model field. The response reports the versioned model that answered. Aliases such as jev-latest move when a new release ships. TypeSafe’s Models page advises: “If you have tuned confidence thresholds against a specific version, pin that version’s ID instead of the alias.”
Use cases reported so far
TypeSafe’s use-case map lists model routing, LLM guardrails, search and re-ranking, semantic code linting, feature extraction, recruiting and customer support. The more useful evidence is what outside developers say they have built. None of the results below were independently replicated for this article.
Classification and routing
The simplest use is the one in the quick start: sort incoming items into buckets. TechCrunch reported two developer comparisons. Pranit Sharma, a software engineer at Vercel, said the company had used an OpenAI model to run a classifier that reviews commands for safety, and that replacing it with Jev produced results “five to 18 times more quickly and with greater accuracy”. Nikhil Mudholkar, CTO of Bryo AI, tested Jev against Gemini on business emails; Gemini was slightly more accurate but 10 to 20 times more expensive in his test, and what interested him more was that Jev “hands back a real probability”.
TypeSafe’s classification using confidence cookbook shows a pattern worth copying. It classified 60 SEC annual reports into 75 industry groups. Filings with confidence of 0.9 or above were right 90 percent of the time; the rest were right 40 percent of the time. Reporting the low-confidence half at the broader division level instead raised their accuracy to 70 percent. Sixty filings is a small sample, and the numbers come from TypeSafe, but the method (fall back to a coarser label when unsure) works with any hierarchical taxonomy.
Agent harnesses
LangChain published a harness guide on 17 September. Its langchain-typesafe package exposes Jev as TypeSafeClassifier and adds two experimental middlewares. ModelRouterMiddleware asks Jev to pick between a fast and a powerful model for each run, based on criteria you write. AutoModeMiddleware uses Jev to “check tool calls for risky decisions” and block them before the tool runs. The post names early projects, including browser agents at Browserbase, a live trading agent and email triage, without giving results.
A second LangChain post on 25 September gave numbers. In a legal document-review graph, Jev answered three questions per page (responsive, contains personal data, possibly privileged), and was 5 to 6 times faster on the classification step than Claude Sonnet in the same graph. Browserbase rebuilt its Stagehand act() step so that Jev picks the action and the page element, and “anything below a 0.7 confidence threshold falls back to an LLM”. Median latency fell from 1.97 seconds to 0.46 seconds in early testing, according to the post.
Evaluation and verification
Jev is being tried as a cheap judge. LangChain reported that in “an early Jev-as-a-judge experiment, its scores barely moved across 100 repeated runs, far less than any LLM judge we tested.” A preprint by Huang and colleagues (arXiv 2609.27607) used Jev to check whether each statement in an AI-generated radiology report is supported by a physician’s reference report. It reached Kendall correlations of 0.573 and 0.398 with expert error counts on two datasets, beat an open natural-language-inference judge under the same setup, and cost “under three cents per hundred report pairs”. The same paper found a specialized local tool, RadMatch, did better on clinically significant errors.
TypeSafe’s SDE cascade cookbook uses Jev as a verifier for structured data extraction: a small LLM extracts, Jev asks one yes-or-no question per field (“is this value absent from the source?”), and if any field’s probability of being wrong exceeds 0.7, the item goes to a reasoning model. TypeSafe labels the resulting cost and quality chart “internal TypeSafe results”.
If you already run LLM-as-a-judge evaluations, the questions in How to compare LLM evaluation frameworks still apply to a Jev judge: agreement with human labels, bias, and whether the scores are stable enough to compare runs.
Moderation and guardrails
TypeSafe’s guardrails cookbook screens every message going into and out of an LLM app with one request: a battery of Noul questions for specific hazards plus a Score for severity. Each hazard has two thresholds. At or above the action threshold it triggers its action (block, or route to support); at or above a lower review threshold the message goes to a person. The cookbook’s point is that the thresholds are policy, set in code, and can change without new model calls. TechCrunch reported that Almeida sees users deploying Jev “to track LLM agent traces and prevent jailbreaks”.
One caution from TypeSafe’s own jaggedness page: jev-1.13 “does not treat [state] as hostile by default”, and content written to steer it “can move the answer”. A guardrail model that can be argued with needs adversarial testing before it is trusted.
Games, trading and real-time demos
The launch post showed Jev playing Doom from a text representation of game state at about 10 queries a second, which TypeSafe says costs about $7 an hour. The Register reported community apps for Doom, Tetris, League of Legends, Settlers of Catan and chess, “where Jev lost to the GLM 5.3 open-weight model but was far cheaper to run”, plus a clothing try-on mockup at $0.0011 a decision and about 620 milliseconds. TypeSafe’s function calling cookbook maps plain-English trading requests onto ten typed functions by turning each closed set of arguments into a Choice.
These demos show speed. They say little about accuracy. The Register quoted YouTuber Mo Bitar: “I know it’s fast. I know it’s cheap, but is it good?”
Jev, an LLM, a reasoning model or a classifier
The table below sets out when each tool fits. It is drawn from TypeSafe’s own statements about what Jev is not for, the jaggedness page, and what the developer reports above describe. It is a starting point, not a verdict; your own test data should decide.
| Situation | Jev | Standard LLM | Reasoning model | Classic trained classifier |
|---|---|---|---|---|
| Output needed | A label, level or yes/no your code branches on | Text, code, JSON, a reply | A worked answer to a hard problem | A label |
| Answer space | Fixed in advance (Choice up to 255 options) | Open | Open | Fixed at training time |
| Labeled training data | Not needed to start | Not needed | Not needed | Needed |
| Per-call latency | TypeSafe claims 70 ms to 500 ms end to end | TypeSafe cites 3 to 329 seconds for frontier models | Varies with reasoning effort | Depends on your hosting |
| Explanation of the answer | None | Can write one | Can write one | None |
| Probabilities | Returned for every option | Not by default | Not by default | Returned; calibratable on your data |
| Arithmetic, counting, dates | Weak, per TypeSafe; do it in code | Mixed | Strongest of the four | Not applicable |
| Multi-step reasoning | Weak on indirection, per TypeSafe | Moderate | Designed for it | Not applicable |
| Changing the task | Edit the question | Edit the prompt | Edit the prompt | Relabel and retrain |
| Where it runs | TypeSafe’s API only | Many providers, some self-hostable | Many providers | Anywhere |
Alex Molas, in an independent critique, put the classifier comparison well: “fine-tuning a BERT requires data. If you don’t have data, Jev is a great alternative.” Once you have a large labeled set for a stable task, a small trained classifier you host yourself may be cheaper, faster and fully under your control. Jev’s case is strongest when the task changes often, labels are scarce, or you need many different judgments about the same input.
Reasoning models remain the right choice when the question itself needs working out. TypeSafe’s jaggedness page says jev-1.13 “may struggle with tasks that require additional levels of indirection”, reads instructions literally, and “does not count reliably”. Its advice for anything numeric is “keep the arithmetic in code”. For what hidden reasoning costs, see The reasoning-model era has a cost problem and Test-time compute, explained.
Patterns for combining Jev with an LLM
The reports converge on three shapes. LangChain’s second post, citing Jaya Gupta, calls the shift one from “frontier by default and optimize later” to “cheap by default, frontier on exception”.
Router in front of the model
Jev reads each incoming request and decides where it goes: deterministic code, a small LLM, a large one, or a person. TypeSafe’s intent routing pattern asks a Choice for intent and a Score for complexity in one call, sends order-status lookups to plain code, sends product and returns questions to specialist LLMs, and sends complex complaints, or any intent with confidence under 0.5, to a human. LangChain’s ModelRouterMiddleware is the same idea applied to model choice.
If you already run an AI gateway, the router decision is one more input to the policy that gateway enforces. What an AI gateway does covers where routing, budgets and fallbacks usually live; a Jev call can sit before the gateway and set a header or model name that the gateway then acts on.
Gate with fallback
Here Jev does the work itself when it is sure, and hands off when it is not. Browserbase’s 0.7 threshold is one example. The SDE cascade is another, run in reverse: the cheap LLM extracts, and Jev decides whether the answer is trustworthy enough to keep. The design choice is where to put the threshold, which is a question about your data, not about the model (see below).
TypeSafe’s confidence-gated routing example, reproduced from its docs, shows thresholds that scale with risk in a voice-banking app:
action = response.answers["intent"]
# Below 0.6 confidence on any action, route to a human
if action.confidence < 0.6:
route_to_support_agent(account_id)
elif action.choice == "check_balance":
# Low stakes. 0.6 confidence is sufficient.
show_balance(account_id)
elif action.choice == "approve_transfer":
if action.confidence > 0.85:
# High stakes, but high confidence. Safe to act automatically.
approve_transfer(account_id)
else:
# High stakes, moderate confidence. Verify intent first.
ask_user_to_confirm("Just to confirm: you would like to approve this transfer, is that correct?")
else:
route_to_support_agent(account_id)
Judge in an evaluation pipeline
In an evaluation pipeline Jev scores outputs that an LLM produced. TypeSafe’s advice for verifier questions, from the SDE cascade cookbook, carries over: make each question “narrow and grounded”, phrase it so that true means something is wrong, ask one question per field or claim, and aggregate with a maximum rather than an average so that “one confident red flag escalates, instead of being averaged into silence”. The radiology preprint used the same decomposition: one support question per statement, in both directions.
Using the probabilities
TypeSafe’s Confidence page proposes three bands: act automatically when confidence is high, proceed with caution (confirm, flag, gather more data) in the middle, and do not act when it is low. It adds that “a confidence threshold is not one number”: a read-only action can run at a lower bar than a destructive one.
Three practices follow from the documentation and the cookbooks.
Make abstention an explicit outcome. TypeSafe’s self-consistency cookbook for Choice questions ran a borderline moderation post through an eight-question rubric 15 times. Jev’s picked labels matched across runs 90.8 percent of the time. When any answer with a top probability under 0.60 was turned into “uncertain” and sent to human review, agreement rose to 99.2 percent, with 25.8 percent of answers going to review. The cookbook is careful to say these figures “measure repeatability only”, not accuracy, and that a Claude Haiku configuration at temperature 0 matched 100 percent of the time with no abstentions.
Use a band, not a line. The companion Noul cookbook found one answer on an insurance claim moving between 0.43 and 0.53 across repeats, crossing a 0.5 threshold. It turned probabilities from 0.30 to 0.70 into an explicit “uncertain” outcome. Any single cut-off will flip on answers that sit near it.
Do not mix thresholds across question types. The jaggedness page shows the same refund question returning 0.22 as a Noul and 0.01 for “yes” as a Choice, and a question and its negation, asked as two Nouls, summing to 1.19. TypeSafe’s advice: “Don’t carry a threshold tuned on a Noul over to a Choice.”
Whatever the bands, log the full probabilities and the versioned model ID with every decision. That record is what lets you retune later and explain a decision to a reviewer.
Checking calibration on your own data
A model is calibrated when its stated probabilities match observed frequencies: of all the answers given 0.8, about 80 percent should be right. TypeSafe says Jev is trained with “reinforcement learning for calibrated decisions” (RLCD), and its AI primer is explicit that calibration rates “describe groups of predictions, not a guarantee about any single answer.”
Outside evidence so far says to verify rather than assume.
- Molas argues that calibration “is not just a property of the model, but also of your data distribution”. Two companies can define spam the same way, have different mixes of spam, and receive identical probabilities from Jev for the same email. His recommendation is to “treat Jev’s outputs as good scores (they rank examples well) rather than good probabilities”.
- Porcedda’s Sys1Cal-v1 preprint (arXiv 2609.35342) built true-or-false questions whose correct probability is known by construction. It says Jev’s calibration claim “is not backed by any public test”, and reports that Choice answers behave as if a third “I don’t know” outcome were being suppressed. Accounting for it raised median soft accuracy on Choice answers from 0.771 to 0.978.
- Rafe and Das (arXiv 2609.24052) coded 195,857 Texas police crash narratives with a 27-question Jev schema. Against 2,416 blinded human judgments, Jev reached an F1 of 0.908. One of two frontier LLMs gained 0.059 over that; the other was indistinguishable. They found that “calibration varies by model rather than by paradigm, so each model must be audited,” and that recalibrating on the same labels cut calibration error by a factor of 3.3.
These are preprints, not peer-reviewed papers. They agree on the practical step, which is standard for any probabilistic classifier:
- Label a sample. Draw a few hundred real inputs from production, not a curated set, and label them by hand. Molas suggests a few hundred can be enough for simple recalibration.
- Plot a reliability diagram. Bin Jev’s probabilities (0 to 0.1, 0.1 to 0.2, and so on) and plot the average predicted probability in each bin against the fraction that were actually correct. scikit-learn’s
calibration_curvedoes this. A calibrated model sits on the diagonal. - Recalibrate if needed. Fit a mapping from Jev’s scores to observed frequencies. The scikit-learn guide describes a sigmoid (Platt) fit as “most effective for small sample sizes”, and says isotonic regression does as well or better “when there is enough data (greater than ~ 1000 samples)”. Guo and colleagues found temperature scaling, a one-parameter version of Platt scaling, “surprisingly effective” for neural networks.
- Pick thresholds from costs. With calibrated numbers, choose each threshold by weighing the cost of a wrong automatic action against the cost of a human review. TypeSafe’s own cookbook says the same: “Choose production thresholds using labeled examples and the cost of incorrect actions and human review.”
- Repeat on every model change. Pin the version. When you move to a new one, rerun the sample before switching.
Do this per question and per primitive. A Noul and a Choice about the same thing are, per TypeSafe, not interchangeable.
Cost arithmetic, with a worked example
TypeSafe’s Models page lists jev-1.13.0 at $0.042 per million input tokens, or $42 per billion. “Output tokens are free.” The response still reports output_tokens, but they are not billed. Question text counts toward input, so every question and option description you add costs a little.
The example below is illustrative. The workload and token counts are assumptions, not measurements, and LLM prices are OpenAI’s standard rates for GPT-6 models as listed on its pricing page on the day of writing: GPT-6 Luna at $0.10 input and $0.50 output per million tokens, GPT-6 Sol at $2.00 and $10.00.
Assumed workload. One million support tickets a month. Each request, state plus three questions, uses 400 input tokens, close to the 392 in TypeSafe’s quick-start response. For the LLMs, assume the same 400 input tokens plus 60 output tokens for a short JSON answer, and no reasoning tokens.
| Setup | Input cost | Output cost | Monthly total |
|---|---|---|---|
| Jev on every ticket | 400M tokens x $0.042/M = $16.80 | $0 | $16.80 |
| GPT-6 Luna on every ticket | 400M x $0.10/M = $40.00 | 60M x $0.50/M = $30.00 | $70.00 |
| GPT-6 Sol on every ticket | 400M x $2.00/M = $800.00 | 60M x $10.00/M = $600.00 | $1,400.00 |
| Jev on every ticket, 15% escalated to Sol | $16.80 + (0.15 x $1,400.00) | $226.80 |
Three points come out of the arithmetic.
The multiple depends on the comparison. Against Sol, Jev is about 83 times cheaper in this example. Against Luna it is about 4 times cheaper, and most of that comes from free output. TypeSafe’s headline figures of “193.6x faster, 444.6x cheaper” come from its own workflow evaluations against frontier models, and the company writes that it expects “these are on the higher end of real world gains.” If your current classifier already runs on a small, cheap LLM, the saving is real but modest, and latency may matter more than price.
The escalation rate dominates a cascade. In the last row, Jev costs $16.80 and the 15 percent that reach Sol cost $210. Halving the escalation rate saves far more than any change to Jev’s price. Calibrating thresholds well is a cost lever, not only a quality one.
Reasoning tokens change the picture. If the Sol calls used reasoning, every hidden reasoning token would be billed as output. At an assumed 500 reasoning tokens per ticket, Sol-on-everything would add 500M x $10/M = $5,000 a month.
Two further caveats. Token counts are not directly comparable across providers, because each uses its own tokenizer; the same text may count differently. And throughput has a ceiling: at the listed 1,200 requests per minute, one million requests take at least 833 minutes, about 14 hours. TypeSafe warns that limits “can change without notice” while it scales, and offers higher limits on enterprise plans.
For controlling spend across several models in one place, see LLM cost management at the gateway.
Limits and risks
Early access. Jev is available to developers as they come off a waitlist. TechCrunch reported that TypeSafe “briefly lost the ability to serve users from its API because demand was so high.” Plan for 429 and 529 errors; the API reference documents both and the SDKs retry by default.
Proprietary and hosted only. Jev runs on TypeSafe’s API. The Models page says the same weights serve every account and Jev “is not fine-tuned or LoRA-adapted with customer data”. You adapt it through the state and questions, not through training. There is no self-hosted option in the documentation.
Lock-in. The request format is TypeSafe’s own. The Register notes that open-source classifiers similar to Jev exist, naming Jeff and Nimble. Neither that report nor this article evaluates them, and moving between them would mean rewriting questions and retuning thresholds. Keeping thresholds, question text and labeled test sets in your own code limits that cost.
“Cannot hallucinate” is a narrow claim. The launch post says Jev “can’t hallucinate” and that it “never makes type errors”. What TypeSafe demonstrates is the second part: the output is always one of the options you defined, so it cannot invent a label or return malformed JSON. It can still pick the wrong option with high confidence. For the broader argument about the word, see Stop calling it hallucination.
Inconsistent vendor numbers. TypeSafe’s figures vary by page. The launch post says 70 ms to 500 ms and 40 to 200 times faster; the press release says “less than 100 milliseconds” and “up to 100 times faster and less expensive”; the workflow evaluations say 193.6 times faster and 444.6 times cheaper. The Register reported TypeSafe’s figure of “as little as 150 ms”. Each is a company claim measured on a different workload.
Language and input. Jev accepts text only. The docs say English “is the primary training language” and other languages are “handled but not equally well”.
What is not public
Several things a buyer would normally check could not be found in TypeSafe’s public materials. Model size and architecture are not disclosed; TechCrunch reported that Almeida is “tight-lipped” and that outside observers suspect Jev is built on an open-weight LLM. The launch post describes its workflow evaluations as scored against the average of GPT-6 Astra and Fable 5.1 rather than human labels, and acknowledges possible bias because the workflows were written by TypeSafe staff. Almeida told TechCrunch the model is trained “exclusively on synthetic data”. The launch post’s FAQ lists questions on public benchmarks and training data, but their answers could not be retrieved from the page when it was fetched for this article.
That leaves independent evidence thin: a few preprints, a handful of developer comparisons reported secondhand, and TypeSafe’s own cookbooks, which are open about being internal. Until more arrives, the safest way to adopt Jev is the one its documentation already recommends: small questions, thresholds you own, a fallback for the uncertain cases, and a labeled sample of your own data to check the numbers against.
Sources
- TypeSafe: Introducing System One Models & Jev (Almeida, 15 September 2026)
- TypeSafe docs: AI primer
- TypeSafe docs: Introduction
- TypeSafe docs: Quick start
- TypeSafe docs: API reference
- TypeSafe docs: State
- TypeSafe docs: Primitives (Questions)
- TypeSafe docs: Confidence
- TypeSafe docs: Models
- TypeSafe docs: Jev 1.13 jaggedness
- TypeSafe docs: Jev with coding agents
- TypeSafe docs: Example use cases
- TypeSafe docs: Confidence-gated routing
- TypeSafe docs: Intent routing
- TypeSafe cookbook: SDE cascade
- TypeSafe cookbook: Guardrails for LLMs
- TypeSafe cookbook: Self-consistency, choices
- TypeSafe cookbook: Self-consistency, nouls
- TypeSafe cookbook: Classification using confidence
- TypeSafe cookbook: Parallel questions
- TypeSafe cookbook: Function calling
- TypeSafe AI emerges from stealth with $40M in funding (Business Wire, 15 September 2026)
- A new kind of AI model from a ChatGPT inventor is thrilling developers (TechCrunch, 18 September 2026)
- Shut up and calculate: Jev's new AI primitives for coders (The Register, 23 September 2026)
- Building a Harness with Jev (LangChain, 17 September 2026)
- Building Prod with Jev and LangGraph (LangChain, 25 September 2026)
- Jev can't be calibrated (Alex Molas, 23 September 2026)
- Jev thinks "I don't know", but doesn't say it: Introducing Sys1Cal-v1 (Porcedda, arXiv 2609.35342)
- Calibrated Decisions at Scale: Converting Police Crash Narratives into Probabilistic Crash Variables (Rafe and Das, arXiv 2609.24052)
- Can Jev Judge Radiology Reports? (Huang et al., arXiv 2609.27607)
- On Calibration of Modern Neural Networks (Guo, Pleiss, Sun and Weinberger, 2017)
- scikit-learn user guide: Probability calibration
- OpenAI API pricing
Questions readers ask
Can Jev replace the LLM in my chatbot or coding agent?
No. TypeSafe's documentation says Jev does not generate text, write code or hold a conversation, and is not a drop-in replacement for the model behind a coding agent. It is designed to make narrow decisions inside an application, often next to an LLM.
How much does Jev cost?
TypeSafe's models page lists jev-1.13.0 at $0.042 per million input tokens ($42 per billion). Output tokens are free. Rate limits are listed as 250,000 tokens per second and 1,200 requests per minute, and the company warns they can change without notice.
Are Jev's probabilities calibrated?
TypeSafe says the model is trained for calibrated decisions and notes that calibration describes groups of predictions, not a guarantee for any single answer. Independent writers and early papers report that calibration varies by task and by question type, and that recalibrating on labeled data reduced calibration error. Measure it on your own data before relying on the numbers.
What confidence threshold should I use?
There is no universal value. TypeSafe's own examples use floors of 0.5 or 0.6 and a higher bar, such as 0.85 or 0.9, for risky actions, and say the right values depend on your domain and should be tuned with your own labeled data.
