Blog · 30 September 2026

TypeSafe's Jev: an AI model that answers in types, not text

TypeSafe's Jev returns typed decisions with calibrated confidence instead of prose. What it is, what the launch numbers do and don't tell you, and where it fits in an AI stack.

AIModelsDeveloper tooling

Most of this year's AI news has been about models that write: longer answers, better code, bigger context windows. The launch that took over our feeds this month goes the other way. TypeSafe's Jev doesn't write anything. You give it some text and a few typed questions, and it hands back values your code can act on, each with a confidence attached.

That sounds small. We think it's one of the more useful ideas to land this year, and one of the easiest to over-apply. Here's what Jev is, what the launch numbers do and don't say, and where we'd reach for it.

01 / 08

What Jev is

Jev is the first model from TypeSafe AI, a San Francisco lab that came out of stealth on 15 September 2026 with a $40 million seed round led by DCVC. Its founders are Diogo Almeida, who worked on RLHF, InstructGPT and ChatGPT at OpenAI, Sasha Sheng and Erik Gafni.

TypeSafe calls Jev a System One model, a nod to Daniel Kahneman's split between fast, intuitive thinking (System 1) and slow, deliberate reasoning (System 2). Chat models are built to reason and explain. Jev is built to make quick, narrow judgements and hand them straight to software.

It's trained differently, too. Instead of RLHF, which rewards answers people like, TypeSafe describes a method it calls Reinforcement Learning for Calibrated Decisions (RLCD): the model's probabilities are optimised against outcomes, so they're meant to reflect real uncertainty.

02 / 08

Three question types

Everything goes through one endpoint, POST /v1/systemone. You send a state, which is the text being judged (a support message, a transaction, a policy), plus a set of named questions. There are three question types:

  • Choice: pick one option from a list. Returns the choice, a probability for every option and a confidence.
  • Score: place the state on an ordered rubric. Returns a score, per-level probabilities and a confidence.
  • Noul: is this statement true? Returns a single probability between 0 and 1.

Here's a trimmed version of the example in TypeSafe's quick start:

{
  "model": "jev-latest",
  "state": "Hi, I've been trying to connect my Stripe account for 3 days and the integration keeps failing. I'm losing sales. Please help ASAP.",
  "questions": {
    "department": {
      "type": "choice",
      "instructions": "Which team should handle this",
      "criteria": {
        "billing": "Payment or subscription issues",
        "technical": "Bugs or integration problems",
        "sales": "Pricing or account questions"
      }
    },
    "is_urgent": {
      "type": "noul",
      "instructions": "The message conveys urgency or time-sensitivity"
    }
  }
}

What comes back is typed values, not a paragraph: "choice": "technical" with a probability for each department, and a number for is_urgent. Every question is evaluated in parallel and in isolation against the same state, so adding questions barely moves the response time.

03 / 08

Decisions, not strings

TypeSafe's own pitch fits in three words: decisions, not strings.

If you've put an LLM inside a product, you've written the glue. A prompt that begs for JSON, a parser, a retry for when the model adds a friendly sentence before the brace, and a fallback for when it invents a fourth category. Jev removes that layer. The answer can't come back in the wrong shape, because the model can only answer inside the shape you gave it.

The more interesting part is the confidence. An LLM will tell you a ticket is "billing" in the same tone whether it's sure or guessing. Jev gives you a number, and TypeSafe's docs build a pattern around it called confidence-gated routing: "The answer tells you what; confidence tells you whether to act." In their example, anything under 0.6 goes to a human, low-stakes actions can run above that, and consequential actions need 0.85 or more before they happen automatically.

That's the part we like most. It turns "how much do we trust the model?" from a feeling into a threshold you can write down, review and change.

04 / 08

The numbers, with salt

TypeSafe says Jev answers in roughly 70 to 500 milliseconds, and that it's up to 193.6x faster and 444.6x cheaper than LLMs on its own System One workflows. Pricing is $0.042 per million input tokens, and output tokens are free.

Two caveats. First, those multipliers come from workflows TypeSafe built, measured against reference models it chose, and TypeSafe itself says they likely represent high-end results. Read them as a direction, not a promise. Second, read "zero hallucinations" carefully. A schema guarantees the answer is one of your options. It doesn't guarantee it's the right one. A model limited to three categories can still confidently pick the wrong one.

Demand hasn't been the problem. Vercel says Jev is the fastest-adopted model in AI Gateway history, used by nearly 13% of its paid teams within 24 hours. TypeSafe paused new signups shortly after opening them, and its docs say rate limits are being adjusted as demand grows. If you can't get a TypeSafe account, Jev is also available through Vercel's AI Gateway as typesafe-ai/jev, and on OpenRouter.

Not everyone is sold. Some well-known developers, Redis creator Salvatore Sanfilippo among them, have argued that the real use cases are narrower than the hype suggests. We half agree, which brings us to where it fits.

05 / 08

Where it fits, and where it doesn't

Jev is a component, not a platform. It shines anywhere a product makes the same small judgement thousands of times:

  • Routing and triage: which queue, which team, which agent, how urgent.
  • Classification and tagging: spam or not, topic, sentiment, policy category.
  • Guardrails: does this output break a rule, does this request need a human.
  • Agent control flow: should the agent take another step, is the task done, did the tests actually pass.

That last one matters to us. Agentic software delivery is full of tiny yes-or-no calls, and today many of them are made by an expensive model writing a paragraph while a regex hopes to find "yes" somewhere in it. A fast, typed, calibrated answer is a much better fit for those gates.

It's the wrong tool for anything that needs generation, reasoning you can read, or long multi-step planning. It's also the wrong tool for high-stakes decisions about people, like hiring, lending or fraud, where you need to explain why and not just how sure. Jev doesn't show its working. For those jobs you want a System Two model, a human, or both.

06 / 08

How we'd wire it in

If you're thinking about trying Jev in a product, this is how we'd approach it:

  1. Start with one decision. Pick a judgement you already make at volume, where you have past examples and a known right answer.
  2. Keep questions atomic. TypeSafe's own guidance is to break complex decisions into small, separate questions and combine the answers in code. Typed output won't rescue a vague question, so choose options and rubrics carefully.
  3. Write thresholds down as policy. Decide, per action, what confidence is needed to act, what triggers a confirmation and what goes to a person. Keep humans on the irreversible calls.
  4. Measure on your own data. Run it alongside what you have today, check how often it's right at each confidence level, and only then let it act on its own.
  5. Keep a fallback. Send low-confidence cases to a person or a larger model. Jev and your LLM work better as a team than as rivals.

07 / 08

The bigger trend

Jev is one launch, but it points at a shift we expect to continue: AI stacks are becoming portfolios. Large reasoning models for planning and drafting. Coding agents for building. And small, fast, typed models for the thousands of decisions in between, each one cheap, predictable and testable like any other function.

For teams that ship software, that's good news. The less of your product depends on parsing prose, the more of it you can test, monitor and trust. We'll keep tracking System One models here as they mature, including whether the calibration holds up outside launch-week demos.

If you're working out where SI fits in your own product, talk to us.