Independent reporting on artificial intelligence.


The Frontier Wire

News

TypeSafe launches Jev, a model that returns decisions instead of text

TypeSafe AI came out of stealth on September 15 with $40 million in seed funding and Jev, a model that answers typed questions about a block of text with probabilities instead of prose. The pricing is unusual, the speed claims are large, and the first independent tests are mixed.

TL;DR

  • TypeSafe AI, a San Francisco lab founded by former OpenAI researcher Diogo Almeida with Erik Gafni and Sasha Sheng, came out of stealth on September 15, 2026 with a $40 million seed round led by DCVC and a waitlisted early-access model called Jev.
  • Jev does not write. It takes a block of text or JSON plus typed questions (Choice, Score or yes/no “Noul”) and returns answers with probabilities. TypeSafe calls this a “System One model.”
  • Pricing is $0.042 per million input tokens, with output free. That structure follows from the design: the answer is a small typed value, not a stream of generated tokens.
  • The speed and cost multiples come from TypeSafe’s own evals, which use two frontier models as the reference answer, and TypeSafe publishes no public benchmark results. The company’s figures also vary from page to page.
  • Early independent tests are mixed. One tester found Jev less accurate than two OpenAI models on a 77-way banking intent task. Another found it more accurate than a local Qwen model on the same dataset, with confidence scores that were strong at the extremes and overconfident in the middle.

A new kind of commercial model went on sale this month, and it cannot write a sentence. On September 15, TypeSafe AI opened early access to Jev, a model that reads text and answers questions about it with typed values and probabilities instead of prose. Where a large language model (LLM) generates tokens one at a time and leaves the calling code to parse them, Jev returns a choice from a list the developer defined, a position on a rubric, or the probability that a statement is true. TypeSafe prices it at $0.042 per million input tokens and charges nothing for output. The company says Jev is two orders of magnitude faster and cheaper than frontier LLMs on the tasks it is built for. This article sets out what launched, who is behind it, how the pricing works, where the headline numbers come from, and what the first independent tests and open questions look like two weeks in. The desk did not test Jev. Every figure below comes from TypeSafe’s own pages or from named third parties.

Jev at a glance

ItemDetailSource
DeveloperTypeSafe AI, San Francisco, founded 2024Business Wire press release
LaunchSeptember 15, 2026, early accessTypeSafe launch post
Current versionjev-1.13.0, served under the aliases jev-latest and jev-previewTypeSafe docs, Models
Model class“System One model”: typed decisions with probabilities, no text generationTypeSafe docs, System One
Question typesChoice (up to 255 options), Score (ordered levels), Noul (yes/no probability)TypeSafe docs, API reference
Training method“Reinforcement Learning for Calibrated Decisions” (RLCD)TypeSafe launch post and AI primer
InputText only: strings, JSON objects, arrays of textTypeSafe docs, Models
Context64k tokens per request; 32k for the state plus the longest questionTypeSafe docs, Models
Price$0.042 per million input tokens ($42 per billion); output freeTypeSafe docs, Models
Rate limits250,000 tokens per second and 1,200 requests per minute, “adjusting dynamically”TypeSafe docs, Models
Fine-tuningNone; the same weights serve every accountTypeSafe docs, Models
AccessWaitlist at typesafe.ai; API, Python and JavaScript SDKs, web playgroundPress release; TypeSafe docs, Quick start
Funding$40 million seed, led by DCVCPress release; DCVC; Wilson Sonsini
Public benchmarksNone published, by choiceTypeSafe launch post FAQ
Parameter count, architectureNot disclosedTypeSafe launch post; TechCrunch

What TypeSafe released

The launch post, signed by Almeida, describes Jev as “a frontier-intelligence function call: unstructured state in, typed probabilistic decisions out.” The company says it built “a new model architecture, parallel sampler for maximum efficiency, and training method we call Reinforcement Learning for Calibrated Decisions (RLCD).”

In practice a developer sends one HTTP request to POST /v1/systemone with two parts. The first is the state: the text or JSON being judged, such as a support ticket, an agent trace or an invoice. The second is a map of questions, each with a type. According to the API reference and introduction, there are three:

  • Choice picks one option from a set the developer defines, with a maximum of 255 options, and returns the chosen option, a probability for each option and a confidence value.
  • Score rates the state against ordered levels the developer describes and returns a probability-weighted score, the distribution across levels and a confidence value.
  • Noul is a yes/no question. It returns a single number between 0 and 1, the probability that the answer is yes, and carries no separate confidence field.

The docs state that every question in a request “is evaluated in parallel and in isolation against the same state in one go,” and that “adding questions barely changes the response time.” A developer can therefore ask a dozen narrow questions about the same ticket in one call and combine the answers in ordinary code. TypeSafe’s own shorthand for this, on its manifesto page, is “smart if-statements.”

The name “System One” borrows from Daniel Kahneman’s Thinking, Fast and Slow, which distinguishes fast, intuitive System 1 judgment from slow, deliberate System 2 reasoning. The System One docs page says “the emphasis is on fast, focused judgments.” Jev itself is named after William Stanley Jevons, the 19th-century economist who argued in The Coal Question (1865) that more efficient steam engines would raise coal consumption rather than cut it. NPR’s Planet Money quotes the line: “It is wholly a confusion of ideas to suppose that the economical use of fuel is equivalent to a diminished consumption.” TypeSafe’s FAQ applies the same logic to AI: “Every order of magnitude drop in the cost of intelligence unlocks orders of magnitude more use cases.”

A longer treatment of the model class, and how it relates to classifiers and constrained decoding, is in the companion explainer, System One models explained.

What Jev is not

TypeSafe’s documentation is direct about the limits. The coding agents page says Jev “is not a drop-in replacement for the LLM behind Claude Code, Cursor, opencode, Copilot” and similar tools, and that it “does not generate text, write code, or hold a conversation.” The home page FAQ adds that “some tasks requiring extended reasoning, such as complex mathematics or chess-like planning, may be better suited to large reasoning models.”

The Jev 1.13 jaggedness page, last reviewed September 17, lists nine known failure modes. Among them: the model reads instructions literally, “does not count reliably,” reads dates “as text, not as ordered quantities,” loses accuracy as the state fills with irrelevant detail, and can be steered by adversarial content in the state because it “does not treat it as hostile by default.” Input is text only; images, audio and video are “not supported (yet).” The Models page says English “is the primary training language and where accuracy is currently best.”

Company, founders and funding

The press release describes TypeSafe as “a frontier AI lab building machine-native, composable AI,” founded in 2024 and headquartered in San Francisco. It names three founders. The team page lists their roles:

  • Diogo Almeida, CEO. The team page says he “co-invented RLHF and InstructGPT, the methods that lead to ChatGPT and GPT4” and was previously at Google Brain. RLHF, reinforcement learning from human feedback, is the post-training method that turned pretrained language models into chat assistants by rewarding responses human raters prefer. In the launch post Almeida writes that at OpenAI he “helped build the methods that made language models useful at following instructions and talking with people.” TechCrunch reports that he left OpenAI two years ago to start the company.
  • Erik Gafni, CTO. Described as a repeat founder (Ravel, multimodal AI for DNA sequencing) and an early employee at Invitae and Freenome.
  • Sasha Sheng, COO. Described as a former research engineer at Meta and FAIR who has published at NeurIPS and ECCV.

The round is $40 million, described as a seed round and led by DCVC, according to the press release, DCVC’s own post and Wilson Sonsini, which advised TypeSafe on the deal. The press release’s boilerplate uses “approximately $40 million.” None of the documents consulted for this article names another investor; TypeSafe’s team page says only that it is “backed by top-tier investors.” No valuation appears in any primary document consulted.

James Hardiman, the DCVC general partner who wrote the firm’s post, said in the press release: “TypeSafe is approaching one of the biggest remaining challenges in AI: turning increasingly capable models into technology that developers can reliably build into products at scale.”

The company’s stated thesis is that the bottleneck in AI adoption is interface, not intelligence. The manifesto argues that current models are “trained to be a helpful, articulate, pleasant assistant,” with the “foreseeable consequence” of “AI that requires humans in the loop instead of running in the background.” Almeida put it more bluntly to TechCrunch: “We have lightning in a bottle, and yet it is not useful.”

Early access and how to get in

Jev is not generally available. The press release says early access “is waitlisted at typesafe.ai,” and the launch post says TypeSafe is “bringing developers off the waitlist as quickly as we can.” Mariya Mansurova, who tested Jev for Towards Data Science, reported that “it took about half a day to receive an invite.”

Once admitted, the quick start offers three routes: a web playground at console.typesafe.ai, the HTTP API with a bearer key, and SDKs for Python (typesafe-sdk) and JavaScript (@typesafe-ai/sdk). Third-party integrations appeared within days. LangChain published a langchain-typesafe package with a TypeSafeClassifier and experimental middleware for model routing and for screening risky tool calls, described in its post Building a Harness with Jev. Langfuse added Jev as a judge for scoring traces, announced September 18. A step-by-step integration guide is in the companion piece, How to use Jev.

Capacity is the main constraint. TechCrunch reported that “the company briefly lost the ability to serve users from its API because demand was so high.” The Models page carries a warning that rate limits “can change without notice while we do, as upcoming large GPU deals land and we let in more users.” Higher limits and zero data retention are offered on enterprise plans through sales. The launch post also notes that the service is “currently based” on the US West Coast.

Pricing: input metered, output free

The Models page lists one price for jev-1.13.0: $42 per billion input tokens, which is $0.042 per million. It states: “Charged per input token. Output tokens are free.” TypeSafe’s launch post frames the contrast with LLMs, which it says charge “from $0.20 to $10” per million input tokens with output “~5x more expensive than input tokens.” The home page claims a “238x lower input price than Claude Fable 5.1.”

Output is not literally zero tokens. The quick start example returns "usage": {"input_tokens": 392, "output_tokens": 65}. The count is reported; it just is not billed.

Why output-free pricing fits a typed model

For a generative model, output is the expensive part of a request. Each output token requires another sequential pass through the model, conditioned on everything before it, so providers charge several times more for output than input. Reasoning models make this sharper, because the hidden reasoning tokens are billed as output too, as covered in The reasoning-model era has a cost problem.

Jev’s output works differently, on TypeSafe’s account. The answer space is fixed before the call, and the launch post says Jev “generates all outputs in a single query” rather than “one token at a time, each conditioned on the last.” The docs say Jev “ingests the state once and evaluates every question against it in parallel.” If that description is accurate, the work of a request scales mainly with how much text the model reads, not with how much it returns. An answer is a label and a handful of probabilities, bounded in size by the schema the developer wrote. Metering output would mean billing for a quantity that tracks the number of options offered rather than the effort the model spent.

The structure has two practical effects. The first is predictability: the cost of a call is known from the size of the state and questions before it is sent, with no risk of a verbose or looping answer inflating the bill. The second is that it rewards packing many questions into one request. TypeSafe’s parallel-questions cookbook, as summarized in its docs index, reports that batching a 13-question briefing into one call was “12.2x cheaper and 10.0x faster with no change in answers.”

Per-token price is not the whole story. Mansurova found that on a 77-way classification task Jev used “about 2× more input tokens and 40× more output tokens” than OpenAI’s models, largely because it returns a probability for every option. The output tokens cost nothing, but she wrote: “I’m not convinced it ends up being dramatically cheaper than Luna for this particular use case.”

Whether the price will hold is a separate question, and TypeSafe’s pages do not answer it the same way. The launch post says: “We can’t prove it isn’t subsidized; we’ll need the long-term to prove the sustainability of our pricing (which we expect to go down, not up).” The home page FAQ says: “We can serve Jev profitably at our current prices.”

The headline claims and how they were measured

TypeSafe makes four headline claims: Jev is much faster, much cheaper, cannot hallucinate, and returns calibrated probabilities. The launch post pairs each with a “Nuance” note, which is unusual in a launch and worth reading in full.

Speed and cost multiples

The numbers vary with the document.

ClaimWhere it appears
“70ms-500ms” end-to-end; “40x-200x faster” for System One queriesTypeSafe launch post
“193.6x Faster, 444.6x Cheaper”TypeSafe home page, sourced to the workflow evals
“Less than 100 milliseconds of latency,” “up to 100 times faster and less expensive”Press release
“Often 100 times faster and less expensive”DCVC post
“20-200x faster,” “40-400x” cheaperAlmeida’s launch post on X, as reproduced by Latent Space
Answers “in as little as 150 ms”The Register, attributing the figure to TypeSafe

The launch post says the 193.6x and 444.6x figures come from its workflow evals and that “we expect that these are on the higher end of real world gains.” It also says of speed: “our published evals are generally run from our laptops on the West Coast.”

Workflow evals

TypeSafe built its own evaluation rather than using public benchmarks. It wrote four workflows (security incidents, agent trace observability, invoice processing and customer service), broke each into narrow Noul, Choice and Score questions joined by code, and ran every model through the same harness. The evals site says: “the reference labels are generated via an average of the responses of GPT-6 Astra and Claude Fable 5.1, both at high thinking.” Accuracy therefore means agreement with those two models, not with human-labeled ground truth.

The launch post lists the caveats itself. The workflows “were made by individuals on our model capabilities team, so some bias could exist.” Using Astra and Fable as the reference “biases answers towards OpenAI and Anthropic’s models.” The LLMs ran through TypeSafe’s own wrapper, which the company says “tends to be slower and more expensive than giving decisions without probabilities.” The evals site does not publish the number of cases per workflow.

The aggregate chart, averaged across the four workflows, shows where Jev sits:

Configuration (workflow harness)Accuracy vs referenceCost per caseTime per case
Jev67.8%$0.00040.4 s
Terra67.9%$0.030410.1 s
Luna66.8%$0.003312.9 s
Sonnet 567.8%$0.117478.1 s
Opus 573.1%$0.176137.8 s
Sol74.1%$0.083623.3 s

Source: evals.typesafe.ai, fetched September 29, 2026. Every model at its provider’s default reasoning setting.

On TypeSafe’s own chart, Jev roughly ties Terra, which the launch post says it found “the most comparable at intelligence to Jev on average,” and trails Sol and Opus 5 by about six points. Against Terra, the published figures work out to about 25 times faster and 76 times cheaper per case. The figures are rounded, and no single comparison model reproduces both 193.6x and 444.6x from them. The nearest matches pair Jev’s time against Sonnet 5 (78.1 s against 0.4 s) and its cost against Opus 5 ($0.1761 against $0.0004). In short, the large multiples appear to come from comparing against the slowest and most expensive models on each axis. The strongest claim the chart supports is that Jev sits far to the cheap, fast end of the accuracy frontier TypeSafe drew. It does not show Jev matching the most accurate models.

“Can’t hallucinate”

The launch post says Jev “can’t hallucinate” and that “the model never makes type errors.” Its supporting chart plots Jev at zero, with the note: “Our number is not empirical. Schema matching is guaranteed, thus we can confidently add 0% into the plots.” The LLM figures on that chart come from OpenRouter, which the company concedes introduces bias.

The guarantee is about form. Jev cannot return an option the developer did not list, produce malformed JSON or answer a yes/no question with a paragraph. It can still choose the wrong option. TypeSafe’s home page FAQ says so plainly: “Jev guarantees the shape of its answers, not that every decision is correct. If you provide a list of categories, it can’t invent a category outside that list, but it can choose the wrong one.” Most readers would call a confident wrong answer a hallucination, which is why the word causes confusion; Stop calling it hallucination sets out why the term covers too many distinct failures. LLMs can also be constrained to a schema with structured outputs, a point several developers raised at launch.

Calibration

A model is calibrated when its stated probabilities match observed frequencies: of all the answers it gives at 0.8, about 80 percent should be right. TypeSafe’s AI primer defines it that way and adds that “these rates describe groups of predictions, not a guarantee about any single answer.” The launch post claims Jev “always communicates confidence and uncertainty with every output” and that “higher confidence means higher accuracy.” TypeSafe has not published a reliability diagram, an expected calibration error or any other calibration measurement for Jev. The claim rests, for now, on the company’s description and on the independent tests below.

The jaggedness page flags one limit on how far the probabilities can be read literally. Asked whether a customer wanted a refund and, separately, whether they wanted something other than a refund, Jev returned 0.72 and 0.47, which sum to 1.19. The docs advise developers not to “hold the model to arithmetic identities between separate questions.”

Independent tests so far

Two weeks after launch, a few developers have published measured tests. None is large, and none was run by this publication.

TesterTaskWhat they found
Mariya Mansurova, Towards Data Science (Sep 21)Banking77, 77 intent classes, against OpenAI’s Terra and Luna with schema-constrained outputJev 79.0% accuracy against 83.9% for Terra and 86.2% for Luna, a gap she calls statistically significant. Jev was “almost 2× faster.” Accuracy “consistently increases for higher-confidence buckets.” With 7 labels instead of 77, results were “roughly on par.”
Nhu Hoang, Towards Data Science (Sep 25)Banking77, all 3,080 test messages, against Qwen3-Coder-Next-80B-A3B run locallyJev 81.1% against 76.4% for Qwen. Median latency 245 ms against 249 ms, not a clean comparison. Qwen returned 53 invalid labels (1.72%); Jev returned none.
Pranit Sharma, Vercel, via TechCrunchClassifier reviewing commands for safety, previously on OpenAI’s LunaResults “five to 18 times more quickly and with greater accuracy,” as reported by TechCrunch.
Nikhil Mudholkar, Bryo AI, via TechCrunchClassifying business emails against GeminiGemini “slightly more accurate, but 10 to 20 times more expensive,” as reported by TechCrunch.

The two Towards Data Science tests used the same dataset and reached different conclusions about accuracy, which says more about setup than about either tester. Mansurova passed labels with no descriptions and compared Jev with two OpenAI models; Hoang passed short descriptions for each category and compared it with an open-weight model running on one GPU.

Hoang’s calibration results are the most detailed public evidence so far. At a confidence of exactly 1.00, Jev was right 97.1 percent of the time across 1,516 messages. Below 0.5 it was right 39 percent of the time against an average stated confidence of 0.41, close to calibrated. The middle was weaker: between 0.7 and 0.9, average confidence was 0.81 but accuracy was 53 percent, and between 0.9 and 0.99 it was 0.95 against 79 percent. “Jev became noticeably overconfident in those ranges,” Hoang wrote. A cascade that sent low-confidence cases to Qwen made results worse at every threshold tested, because Qwen got only 57 percent of the hard cases right. Hoang’s conclusion: “neither a clean output format nor a high confidence score should be trusted on its own.”

Every’s head of evals, Mike Taylor, published a review on launch day; most of it sits behind a paywall. Its public opening concludes: “Just how well it gets the job done is still an open question.”

Developer and industry reaction

The launch drew wide attention. The Hacker News thread for the launch post had 1,987 points and 520 comments when fetched for this article, and Latent Space’s AINews led with it on launch day. Almeida’s post on X announcing Jev showed 4.21 million views in the embed Latent Space reproduced.

Much of the enthusiasm came from developers who already use LLMs as narrow classifiers. “After much fumbling around with prompts and evals, this is exactly how I am using LLMs in production,” one Hacker News commenter wrote. Armin Ronacher, CTO of Earendil, told TechCrunch that Jev “delegates the hallucination problem a little bit to the user,” and that a 50 percent answer might be treated as “a coin toss” while a 95 percent answer could be acted on. He also pointed to model routing as a use case and said he expects competitors to follow. The Register reported that FPV Ventures partner Nikunj Kothari set up a site, Jevable, to collect prototype apps, and quoted Andrej Karpathy writing on X that Jev “revealed latent demand […] that was under-invested into because of a race to higher intelligence.”

Skepticism focused on the claims, not the idea. On Hacker News, one commenter called the speed comparison “misleading” because a generative model can output code while Jev cannot, and added: “‘can’t hallucinate’ seems wrong? Sure, it can’t emit an invalid type, but it can still emit a completely wrong valid value.” Another wrote that “all claims just sound like marketing terms,” calling the “70-500ms vs 3-329 seconds” comparison “apples-to-oranges.” A third noted that encoder classifiers already offered probabilistic outputs with no generation, and asked whether Jev’s advance is mainly letting developers define the output space without training data. The Register quoted YouTuber Mo Bitar: “I know it’s fast. I know it’s cheap, but is it good?”

Open questions

Benchmarks. TypeSafe says it “deliberately chose not to publish performance against public benchmarks” and plans “to only have one-off evals when we make product updates.” It recommends that users build their own evals, which is sound advice for any model and covered in How to compare LLM evaluation frameworks. It also means there is no shared yardstick for comparing Jev with anything else, and the only company-published accuracy figures measure agreement with two other models.

Size and architecture. TypeSafe discloses neither. Its FAQ says only that “Jev is neither small nor an LLM.” TechCrunch describes it as “transformer-based” and reports that outside observers suspect it is built on top of an open-weight LLM. The Register’s report states that “Jev is an LLM.” TypeSafe’s AI primer shows a figure whose caption describes pretrained language models branching into RLHF, RLVR and “an emphasized RLCD decision-model path,” which suggests RLCD is applied after pretraining, but the company has not said what it starts from. There is no model card; Reading a model card like an auditor lists what one would normally contain.

Training data and method. TypeSafe says “we make all the data ourselves” and does not train on customer data. Almeida told TechCrunch Jev is trained “exclusively on synthetic data.” RLCD has no paper or technical report. Asked for details, the launch FAQ answers: “if you want to find out more, we’d have to hire you.”

Calibration evidence. The core selling point has no published measurement from TypeSafe. Hoang’s test found good calibration at the extremes and overconfidence in the middle on one dataset. Calibration can differ by task, language and question wording, and the docs advise teams that have tuned confidence thresholds to pin a specific version rather than follow an alias.

Availability and licensing. Jev is API-only, waitlisted, served from the US West Coast, and subject to rate limits that can change without notice. No weights are published and there is no fine-tuning. Zero data retention is enterprise-only. One Hacker News commenter asked for access through established platforms such as OpenRouter or AWS Bedrock to avoid a new vendor review; TypeSafe has not announced any such distribution.

Lock-in. The request shape is TypeSafe’s own, and no other provider serves a compatible endpoint. The jev-latest alias moves when a new release ships, so “the answers behind it can change without a change on your side,” the docs warn, recommending that teams pin a version ID. TypeSafe has published an MIT-licensed System One Adapter, a “drop-in TypeSafeClient replacement backed by LLM APIs” from OpenAI, Anthropic or Google. That gives code written against Jev an exit path to an LLM, at LLM prices and speeds.

Security. The jaggedness page says content designed to steer the model “can move the answer.” That matters for the guardrail and jailbreak-detection uses TypeSafe and LangChain promote, where the input is by definition potentially hostile.

What to watch next

  • General availability and capacity. When the waitlist ends, and whether the rate limits settle as the GPU deals the docs mention come online.
  • New versions. The jev-preview alias exists but points at the same model as jev-latest for now. The jaggedness page says many failure modes “will be fixed in later versions.”
  • New modalities. TechCrunch reports that TypeSafe “will be building more versions of the model, in new modalities.” The docs say image, audio and video input are “not supported (yet).”
  • Evidence. Whether TypeSafe publishes calibration measurements, case counts for its workflow evals, or a technical report on RLCD, and whether larger independent studies appear.
  • Price. Whether the $0.042 rate holds as usage grows. The company has said it expects prices to fall.
  • Competition. Ronacher told TechCrunch he expects rivals “now that its utility is apparent.” Whether large model providers expose comparable decision endpoints, or whether gateways and routers start treating typed decision models as a separate tier, will shape how much of this market TypeSafe keeps.

For now, Jev offers developers something no major provider sold before: a model priced and built only for the narrow judgments software makes thousands of times a day. How accurate those judgments are, and how far the probabilities can be trusted, is being measured by users rather than by the company.

Sources

  1. Introducing System One Models & Jev (Diogo Almeida, TypeSafe AI, September 15, 2026)
  2. TypeSafe AI home page and FAQ
  3. TypeSafe AI team page
  4. TypeSafe AI manifesto: Build Prod, Not God
  5. TypeSafe docs: Models (jev-1.13.0 pricing, limits, aliases)
  6. TypeSafe docs: System One
  7. TypeSafe docs: Introduction
  8. TypeSafe docs: Quick start
  9. TypeSafe docs: AI primer (RLCD)
  10. TypeSafe docs: Confidence
  11. TypeSafe docs: Jev 1.13 jaggedness
  12. TypeSafe docs: Jev with coding agents
  13. TypeSafe docs: API reference
  14. TypeSafe docs: Legal
  15. TypeSafe workflow evals
  16. System One Adapter (TypeSafe, MIT licence)
  17. TypeSafe AI Emerges From Stealth With $40M in Funding (Business Wire press release, September 15, 2026)
  18. TypeSafe emerges from stealth with a new way of doing AI (James Hardiman, DCVC)
  19. Wilson Sonsini advises TypeSafe AI on $40 million seed round
  20. A new kind of AI model from a ChatGPT inventor is thrilling developers (Tim Fernholz, TechCrunch, September 18, 2026)
  21. Shut up and calculate: Jev's new AI primitives for coders (Joab Jackson, The Register, September 23, 2026)
  22. A New Kind of Model for AI Decision-Making? (Mariya Mansurova, Towards Data Science, September 21, 2026)
  23. Jev vs. LLMs: When AI Moves from Generation to Decision-Making (Nhu Hoang, Towards Data Science, September 25, 2026)
  24. Building a Harness with Jev (LangChain blog)
  25. Using TypeSafe's Jev for evals (Langfuse, September 18, 2026)
  26. Mini-Vibe Check: TypeSafe's Jev (Mike Taylor, Every, September 15, 2026)
  27. [AINews] Jev: a System One Model (Latent Space, September 15, 2026)
  28. Introducing System One Models and Jev (Hacker News discussion)
  29. Why the AI world is suddenly obsessed with Jevons paradox (NPR Planet Money, 2025)

Questions readers ask

What is Jev?

Jev is a model from TypeSafe AI, released in early access on September 15, 2026. It does not generate text. A developer sends a block of text or JSON (the state) plus typed questions, and Jev returns a choice, a score or a yes/no probability for each, along with probability distributions and, for Choice and Score questions, a confidence value. TypeSafe calls this class of model a System One model.

How much does Jev cost?

TypeSafe's models page lists jev-1.13.0 at $0.042 per million input tokens, or $42 per billion. Output tokens are free. Responses still report an output token count, but it is not billed.

Can Jev hallucinate?

Jev cannot return a value outside the answer set the developer defines, so it cannot invent a category or produce a malformed response. It can still pick the wrong option. TypeSafe's own FAQ says Jev "guarantees the shape of its answers, not that every decision is correct."

How do I get access to Jev?

Access is through a waitlist at typesafe.ai. Once admitted, developers get an API key from the TypeSafe console and call POST /v1/systemone, or use the Python or JavaScript SDKs. One independent tester reported receiving an invite in about half a day.

Has Jev been independently benchmarked?

Not on public benchmarks. TypeSafe says it deliberately chose not to publish public benchmark results. A handful of developers have published their own tests, with mixed accuracy results; see the independent tests section of this article.

More news