JA

What is Jev? TypeSafe's System One Model for software decisions, not text

Introduction

In September 2026, TypeSafe AI released its first public model, Jev. Jev isn’t an LLM for chat or code generation. It’s a model built to return the “decisions” an application needs internally — classification, routing, scoring, verification — as typed output.

This article draws on TypeSafe’s official documentation, its official blog, and other publicly available information to cover:

  • What TypeSafe AI and Jev are
  • How Jev’s design philosophy differs from generative LLMs
  • Where it fits inside a system
  • The state of the official Playground
  • jeff, a local-compatible implementation
  • How to think about production adoption today

The short version: I see Jev not as a replacement for generative LLMs, but as an option for handling small, repeated decisions inside an application as typed output. That said, it just launched, it’s closed-weight with no self-hosting option, and I’d think carefully before committing to it in production given the dependency on a single vendor.

As of this writing, TypeSafe’s official Console has paused new signups due to a surge in demand, so I haven’t been able to test the Playground hands-on. Once official access reopens, I plan to add a comparison between the real Jev and the local-compatible implementation.

Here is what this post covers:

What is TypeSafe AI

TypeSafe AI is a San Francisco-based AI company founded in 2024. On September 15, 2026, it emerged from stealth, announcing roughly $40 million in DCVC-led seed funding alongside its first public model, Jev.

What TypeSafe AI is building toward isn’t an AI that talks back to a human in a chat window — it’s AI your software can call directly, compose, and constrain. The company calls this composable AI.

Jev is the first model to ship under that vision.

What Jev actually does

Instead of generating free-form text, Jev returns typed decisions that an application can use directly.

Given an inquiry as input, for example, it can return several judgments at once:

Input:
"This is the third time I've contacted support and it's still
 not resolved. I want to talk to a person right now."

Decisions to make:
- Should this escalate to a human agent?
- Which team should handle it?
- How urgent is it?
- Does it contain threatening or hostile language?
- Is it about a refund or billing?

A traditional LLM can approximate this with a JSON Schema. But Jev isn’t built around “generate text, then parse that text as JSON” — it’s designed to return the decision directly from the start.

TypeSafe’s SDK is built around three question types:

TypeReturnsExample
NoulThe probability that a statement is true (0-1)“Is the customer explicitly asking for a human?”
ChoiceOne selection from a defined set of options“Is this billing, technical, or account?”
ScoreA value on an ordered scale“Is urgency low, medium, or high?”

Noul isn’t a standard term — it’s the name TypeSafe gave to this “probability that a statement is true” type.

How it differs from LLMs

The typical way to use an LLM today

If you ask a general-purpose LLM to classify a support ticket, you’d typically write a prompt like this:

Classify the following inquiry.

Inquiry:
This is the third time I've contacted support and it's still
not resolved. I want to talk to a person right now.

Return only the following JSON.

{
  "requires_human": true,
  "team": "technical",
  "urgency": "high"
}

This approach is flexible, but in practice it tends to run into:

  • Text mixed in with the JSON, breaking the parser
  • Values outside the enum you specified
  • Malformed JSON
  • Regenerating the entire output just to change one field
  • Cost and latency growing with output tokens
  • No clear confidence value, making threshold-based branching hard to design

Structured Outputs and function calling have improved a lot of this. Still, the model’s core job remains “predict the next token.”

Jev’s design philosophy

Jev is positioned not as a text-generation model, but as one that computes answers to defined questions against an input state.

state (the information being judged)
    +
questions (typed questions you've defined)
    ↓
Jev / System One Model
    ↓
typed answers (structured decisions, probabilities, confidence)

The name “System One” comes from Daniel Kahneman’s Thinking, Fast and Slow — the fast, intuitive, automatic mode of thought he calls System 1. Jev isn’t aiming for complex reasoning or generating explanations; it’s aiming to be the layer that handles the huge volume of small decisions inside software, fast.

Comparison

Typical generative LLMJev / System One Model
Main purposeConversation, writing, summarization, code generation, complex reasoningClassification, routing, verification, scoring, conditional logic
Basic outputFree-form textTyped decisions, probabilities, confidence
ImplementationGenerate JSON, then parse and validate itStructured output your code can use directly
Good fit forInterpreting ambiguous requests, explanations, writing, reasoningHigh-volume, repetitive, low-latency judgments
Common failure modesHallucination, broken JSON, out-of-enum values, verbose outputMisclassification, low confidence, poorly designed questions
How you improve itPrompting, RAG, model swaps, fine-tuningRefining instructions/criteria, breaking down questions, thresholds, eval data
How you evaluate itHumans reading and judging output qualityComparing against ground-truth labels, which is easy to automate

As the table shows, Jev isn’t a replacement for LLMs. Work like “write a courteous reply to a customer,” “reason across several documents to find the root cause of an incident,” or “draft a design doc” still needs a generative LLM.

But for the layer of decisions — “which process should this route to,” “should this risky action go to a human for approval,” “does this input meet a given condition” — a model like Jev is a better fit.

Where it fits inside a system

The point of Jev isn’t to swap out your existing LLM wholesale — it’s to hand off the “small, repeated, structured decisions” already scattered through an application. Here’s what calling the SDK looks like, and why logging and evaluation matter once it’s running.

What calling the SDK looks like

Here’s a Python example that classifies a support inquiry.

from typesafe_sdk import Choice, Noul, Score, TypeSafeClient

client = TypeSafeClient(api_key="YOUR_API_KEY")

result = client.system_one(
    "This is the third time I've contacted support and it's still "
    "not resolved. I want to talk to a person right now.",
    {
        "human_escalation": Noul(
            instructions="Is the customer clearly asking to speak with a human agent?"
        ),
        "team": Choice(
            instructions="Select the team that should handle this inquiry.",
            criteria={
                "billing": "Billing, refunds, payments, or duplicate charges",
                "technical": "Outages, bugs, features, or how-to questions",
                "account": "Login, authentication, permissions, or account settings",
                "other": "Doesn't clearly fit any of the above",
            },
        ),
        "urgency": Score(
            instructions="Rate how urgent this inquiry is.",
            criteria=[
                "low: a routine inquiry with no expectation of immediate response",
                "medium: the customer is frustrated but there's no outage or hard deadline",
                "high: outage, financial loss, strong demand for immediate action, or a hard deadline",
            ],
        ),
    },
)

print(result.nouls["human_escalation"].noul)
print(result.choices["team"].choice)
print(result.scores["urgency"].score)

What a human designs here is the “shape of the question”:

  • What to have the model decide
  • What options to offer
  • How to define each option
  • What confidence threshold is safe to auto-act on
  • Who to escalate to when confidence is low

The model, meanwhile, is the one actually deciding whether human_escalation is true, whether team is technical, and whether urgency is high.

Logging pays off later

Because Jev’s output is structured, it drops straight into your application logs as-is. Compare that output against human-confirmed ground truth later, and you get continuous, measurable accuracy tracking.

{
  "request_id": "req_01H...",
  "model": "jev-<pinned-version>",
  "question_set_version": "support-routing-v3",
  "state_hash": "sha256:...",
  "answers": {
    "human_escalation": {
      "value": true,
      "probability": 0.97
    },
    "team": {
      "value": "technical",
      "confidence": 0.81
    },
    "urgency": {
      "value": 1.8,
      "confidence": 0.76
    }
  },
  "final_action": "handoff_to_human",
  "human_label": {
    "human_escalation": true,
    "team": "technical",
    "urgency": "high"
  }
}

Here, urgency’s value is recorded on a 0-2 scale corresponding to low/medium/high.

At minimum, your logs should capture:

  • The exact pinned model version used for inference
  • The question-set version, including instructions and criteria
  • The raw input text, or a hash / safely extracted features if you need to avoid storing PII
  • Not just the selected value, but confidence or the full probability distribution
  • The final action your code took
  • The ground-truth label a human confirmed afterward

With logs and ground-truth labels in hand, you can track model performance and operational risk along several axes:

  • Accuracy
  • Precision, recall, and F1-score
  • The confusion matrix, to spot which team or category gets misclassified most
  • Auto-execution and mis-execution rates at each confidence threshold
  • Calibration — checking whether a prediction at 0.9 confidence is actually right about 90% of the time
  • Regression testing — confirming that a model or question-set update hasn’t degraded performance on the same eval set

Thresholds, calibration, and regression testing are beyond this article’s scope, so I’ll leave it at an overview. TypeSafe’s own recommendation is to route by confidence — auto-execute, confirm, or escalate to a human — and tune thresholds against your own data and the risk of your use case. For more, see TypeSafe: Confidence and Confidence-gated routing.

Worth noting: a confidence of 0.9 doesn’t guarantee that single prediction has a 90% chance of being correct. Whether 0.9 actually tracks real-world accuracy is something you can only determine by collecting enough examples for your specific use case and checking the actual accuracy rate within each confidence band.

The state of the official Playground

TypeSafe offers a Playground in its official Console, where you can type in text and try out Noul, Choice, and Score questions. The official Quick Start also points to the Playground as the entry point.

As of this writing, though, demand has spiked, and the sign-in screen showed this message in my environment:

Whoops, we're full

The URL included signups_disabled, confirming new account creation is paused.

Once I can get in, here’s what I plan to test:

  • Classification accuracy on Japanese-language inquiries
  • Output stability across repeated identical inputs
  • How accuracy shifts with the granularity of instructions and criteria
  • Latency when batching multiple questions into a single request
  • Output differences between the real Jev and the local-compatible implementation
  • How to design thresholds for routing low-confidence cases to human review
  • Regression testing against a pinned model version

Trying it locally with jeff

The real Jev doesn’t publish its weights, so there’s no local execution or self-hosting option. You can’t fire up a real Jev with something like ollama run jev.

There is, however, a local-compatible implementation on GitHub that matches Jev’s API shape. The notable one is logan-markewich/jeff.

  • GitHub: logan-markewich/jeff
  • What it does: self-hosts a Jev-compatible /v1/systemone API
  • Backend: the open model GLiFormer (knowledgator/gliformer-large-v1)
  • Notable: works with the official typesafe-sdk as-is, so you can try Noul, Choice, and Score locally
  • Caveat: accuracy, confidence, and latency don’t match the real Jev
client = TypeSafeClient(
    api_key="devkey",
    base_url="http://localhost:8000",
)

Point base_url (or the TYPESAFE_BASE_URL environment variable) at your local server, and your application code can stay exactly the same while you work out question design, data shape, logging, and branch logic against a local instance.

jeff bills itself as a drop-in replacement, but the model running underneath it is a different one entirely. I’d treat it less as a substitute for Jev and more as a development and learning environment for trying out System One-style API design. If you want to evaluate the real model’s quality, confidence calibration, or operational SLA, you’ll need to wait for official API access and compare against the same eval dataset.

How to think about adoption

Right now, I think it’s too early to put Jev at the core of a production architecture. The reasons are straightforward:

  • It just launched, with little long-term operational track record yet
  • The Console’s signup pause is itself a signal worth watching — how well it scales with demand isn’t proven yet
  • The real model is closed-weight; you can’t self-host it
  • It’s not a fit as-is for systems handling sensitive data or PII you can’t send to an external API
  • Its decisions are probabilistic — it doesn’t fully replace business rules or human review
  • If you can’t fully trust your input data, you need to design for prompt injection and similar risks

That said, a PoC is worth doing if any of the following apply:

  • You’re running high volumes of classification, routing, or relevance judgments over and over
  • You’re repeatedly parsing and retrying a generative LLM’s JSON output
  • You need low latency and confidence-based branching
  • You can put together a small eval set with human-confirmed ground truth

If you do adopt it, abstract away the dependency on the official API so you can swap in a different implementation later. For example, put a layer like this in your application:

DecisionProvider
  ├─ TypeSafeJevProvider
  ├─ LocalJeffProvider
  └─ LlmStructuredOutputProvider

That way, if Jev’s offering, pricing, accuracy, or usage limits change, you can switch implementations without rewriting your business logic.

Summary

TypeSafe AI’s Jev isn’t a text-generation LLM — it’s a System One Model built to return decisions inside your software as typed output.

What’s interesting isn’t just Jev’s performance as a single model. It’s the architecture behind it: instead of handing everything to a generative LLM, it splits the work by role.

Decision model (System One Model):
classification, routing, verification, scoring, guardrails

Generative LLM:
explanation, summarization, conversation, writing, complex reasoning

That split has real potential to improve cost, latency, reliability, and auditability when you’re putting AI workflows into production.

I’d start with a local-compatible implementation like jeff to work out question design, logging, an eval dataset, and threshold-based escalation. Once the real Jev’s Playground or API is available, compare accuracy, calibration, speed, and cost on the same data. Taken in that order, you can pull the System One design pattern into your own project without getting carried away by any one vendor’s hype.

References

TypeSafe official

Local-compatible implementation

Evaluation metrics and operations