Jev explained: the AI model that can't chat, what it costs and what it's for
Jev is the AI model developers have been talking about since mid-September 2026, and the odd thing about it is what it doesn’t do. It can’t chat, write an email, draw a picture or write code. It makes decisions. Here is what that means, what it costs, and how much of the hype is backed by evidence.
Who makes Jev
Jev comes from TypeSafe AI, a San Francisco lab founded in 2024 by Diogo Almeida, Erik Gafni and Sasha Sheng. Almeida spent about four years at OpenAI and worked on RLHF, InstructGPT, ChatGPT and GPT-4, which is why the press calls Jev a model “from a ChatGPT inventor”. TypeSafe opened early access on September 15, 2026, with a waitlist.
What Jev actually does
TypeSafe calls Jev a “System One model”. A normal language model writes text, and your code then has to read that text and work out what it means. Jev skips the text. You send it an input (text or JSON; each request has a 64,000-token budget, of which 32,000 can go to the input and question) plus the answer format you want, and it returns a typed answer your program can use directly:
- Choose between options you define (up to 255 of them), with a probability for each.
- Score something on a scale you define.
- Yes or no, as a probability between 0 and 1.
Choices and scores come with a confidence score; a yes/no answer is itself a probability. Typical uses are routing support tickets, moderating content, classifying documents or deciding which tool an agent should call next: small, repeated decisions where a chatbot is slow and expensive.
Developers use it through a REST API with Python and JavaScript SDKs. The current model is jev-1.13.0.
What Jev costs
| Item | Price |
|---|---|
| Input | $0.042 per million tokens ($42 per billion) |
| Output | Free |
That is far below any chat model we track. For comparison, the cheapest models on our price list start at $0.10 per million input tokens, and frontier models charge several dollars. TypeSafe has not published a free tier or monthly plans. Besides TypeSafe’s own early access, Jev 1.13 is available in beta on OpenRouter at the same price.
The catch is that a price per token doesn’t tell you the whole story, because Jev does a narrower job. It replaces the deciding part of a workflow, not the writing part.
Is it really faster and cheaper?
The headline numbers come from TypeSafe itself:
- Maker’s claim: 40 to 200 times faster than frontier language models on “System One tasks”.
- Maker’s claim: up to 444.6 times cheaper on four internal workflows.
TypeSafe says openly that these workflows were built by its own team and that the results sit at the higher end of what to expect. It also says it will not publish a standard benchmark table, arguing that benchmarks get gamed.
Independent evidence so far is thin but useful:
- TechCrunch reports developers’ early experiences: Vercel said Jev ran five to 18 times faster than the OpenAI model it replaced, with better accuracy; another developer found Gemini slightly more accurate but 10 to 20 times more expensive. These are anecdotes, not benchmarks.
- PriorBench, a pre-registered test by an anonymous researcher, ran about 5,700 calls and published the raw data. Its author calls it a “wide, shallow pass”, so treat it as informal.
- Check Point tested security and found Jev’s verdicts could be flipped with prompt injection in every setup it tried, even when Jev was told to ignore such instructions.
When we checked on September 28, 2026, we found no Jev entry on LMArena or Artificial Analysis.
The limitations
- “Can’t hallucinate” is narrower than it sounds. Jev always answers with one of your options, but a well-formed answer isn’t necessarily a correct one. The Register’s question is the right one: it’s fast and cheap, but is it good?
- You design the options. Jev can’t handle a situation you didn’t anticipate in your schema.
- Confidence needs interpreting. A 50% answer might be a coin toss; you decide the thresholds.
- It’s a black box. There is no paper, and the architecture and weights are not published.
- It can be manipulated. Check Point’s results mean you shouldn’t let untrusted text drive high-stakes decisions without other checks.
Who should try it
If you run an app that makes the same kind of decision thousands of times a day (classify this, route that, approve or flag), Jev is worth testing against the language model you use now. Measure accuracy on your own examples before switching. If you need text, code or a conversation, Jev is not the tool; see our price comparison and AI apps list instead.
What’s your experience with it? Tell others in Discussions.
Sources
- TypeSafe AI: Introducing System One models and Jev
- TypeSafe docs: models
- TypeSafe docs: introduction
- TypeSafe AI: Antibenchmaxxing
- Wikipedia: Jev (AI model)
- TechCrunch: A new kind of AI model from a ChatGPT inventor is thrilling developers
- The Register: Shut up and calculate: Jev's new AI primitives for coders
- The Register: TypeSafe AI debuts model for machines
- Bloomberg: Jev, an AI model that can't chat, takes on bigger rivals
- Check Point: Jev is not a language model, but it breaks like one
- OpenRouter: Jev 1.13
- PriorBench: independent Jev test (GitHub)