What is Jev? TypeSafe's System One Model for software decisions, not text
Introduction
In September 2026, TypeSafe AI released its first public model, Jev. Jev isn’t an LLM for chat or code generation. It’s a model built to return the “decisions” an application needs internally — classification, routing, scoring, verification — as typed output.
This article draws on TypeSafe’s official documentation, its official blog, and other publicly available information to cover:
- What TypeSafe AI and Jev are
- How Jev’s design philosophy differs from generative LLMs
- Where it fits inside a system
- The state of the official Playground
jeff, a local-compatible implementation- How to think about production adoption today
The short version: I see Jev not as a replacement for generative LLMs, but as an option for handling small, repeated decisions inside an application as typed output. That said, it just launched, it’s closed-weight with no self-hosting option, and I’d think carefully before committing to it in production given the dependency on a single vendor.
As of this writing, TypeSafe’s official Console has paused new signups due to a surge in demand, so I haven’t been able to test the Playground hands-on. Once official access reopens, I plan to add a comparison between the real Jev and the local-compatible implementation.
Here is what this post covers:
- What is TypeSafe AI
- What Jev actually does
- How it differs from LLMs
- Where it fits inside a system
- The state of the official Playground
- Trying it locally with
jeff - How to think about adoption
- Summary
- References
What is TypeSafe AI
TypeSafe AI is a San Francisco-based AI company founded in 2024. On September 15, 2026, it emerged from stealth, announcing roughly $40 million in DCVC-led seed funding alongside its first public model, Jev.
What TypeSafe AI is building toward isn’t an AI that talks back to a human in a chat window — it’s AI your software can call directly, compose, and constrain. The company calls this composable AI.
Jev is the first model to ship under that vision.
What Jev actually does
Instead of generating free-form text, Jev returns typed decisions that an application can use directly.
Given an inquiry as input, for example, it can return several judgments at once:
Input:
"This is the third time I've contacted support and it's still
not resolved. I want to talk to a person right now."
Decisions to make:
- Should this escalate to a human agent?
- Which team should handle it?
- How urgent is it?
- Does it contain threatening or hostile language?
- Is it about a refund or billing?
A traditional LLM can approximate this with a JSON Schema. But Jev isn’t built around “generate text, then parse that text as JSON” — it’s designed to return the decision directly from the start.
TypeSafe’s SDK is built around three question types:
| Type | Returns | Example |
|---|---|---|
Noul | The probability that a statement is true (0-1) | “Is the customer explicitly asking for a human?” |
Choice | One selection from a defined set of options | “Is this billing, technical, or account?” |
Score | A value on an ordered scale | “Is urgency low, medium, or high?” |
Noul isn’t a standard term — it’s the name TypeSafe gave to this “probability that a statement is true” type.
How it differs from LLMs
The typical way to use an LLM today
If you ask a general-purpose LLM to classify a support ticket, you’d typically write a prompt like this:
Classify the following inquiry.
Inquiry:
This is the third time I've contacted support and it's still
not resolved. I want to talk to a person right now.
Return only the following JSON.
{
"requires_human": true,
"team": "technical",
"urgency": "high"
}
This approach is flexible, but in practice it tends to run into:
- Text mixed in with the JSON, breaking the parser
- Values outside the enum you specified
- Malformed JSON
- Regenerating the entire output just to change one field
- Cost and latency growing with output tokens
- No clear confidence value, making threshold-based branching hard to design
Structured Outputs and function calling have improved a lot of this. Still, the model’s core job remains “predict the next token.”
Jev’s design philosophy
Jev is positioned not as a text-generation model, but as one that computes answers to defined questions against an input state.
state (the information being judged)
+
questions (typed questions you've defined)
↓
Jev / System One Model
↓
typed answers (structured decisions, probabilities, confidence)
The name “System One” comes from Daniel Kahneman’s Thinking, Fast and Slow — the fast, intuitive, automatic mode of thought he calls System 1. Jev isn’t aiming for complex reasoning or generating explanations; it’s aiming to be the layer that handles the huge volume of small decisions inside software, fast.
Comparison
| Typical generative LLM | Jev / System One Model | |
|---|---|---|
| Main purpose | Conversation, writing, summarization, code generation, complex reasoning | Classification, routing, verification, scoring, conditional logic |
| Basic output | Free-form text | Typed decisions, probabilities, confidence |
| Implementation | Generate JSON, then parse and validate it | Structured output your code can use directly |
| Good fit for | Interpreting ambiguous requests, explanations, writing, reasoning | High-volume, repetitive, low-latency judgments |
| Common failure modes | Hallucination, broken JSON, out-of-enum values, verbose output | Misclassification, low confidence, poorly designed questions |
| How you improve it | Prompting, RAG, model swaps, fine-tuning | Refining instructions/criteria, breaking down questions, thresholds, eval data |
| How you evaluate it | Humans reading and judging output quality | Comparing against ground-truth labels, which is easy to automate |
As the table shows, Jev isn’t a replacement for LLMs. Work like “write a courteous reply to a customer,” “reason across several documents to find the root cause of an incident,” or “draft a design doc” still needs a generative LLM.
But for the layer of decisions — “which process should this route to,” “should this risky action go to a human for approval,” “does this input meet a given condition” — a model like Jev is a better fit.
Where it fits inside a system
The point of Jev isn’t to swap out your existing LLM wholesale — it’s to hand off the “small, repeated, structured decisions” already scattered through an application. Here’s what calling the SDK looks like, and why logging and evaluation matter once it’s running.
What calling the SDK looks like
Here’s a Python example that classifies a support inquiry.
from typesafe_sdk import Choice, Noul, Score, TypeSafeClient
client = TypeSafeClient(api_key="YOUR_API_KEY")
result = client.system_one(
"This is the third time I've contacted support and it's still "
"not resolved. I want to talk to a person right now.",
{
"human_escalation": Noul(
instructions="Is the customer clearly asking to speak with a human agent?"
),
"team": Choice(
instructions="Select the team that should handle this inquiry.",
criteria={
"billing": "Billing, refunds, payments, or duplicate charges",
"technical": "Outages, bugs, features, or how-to questions",
"account": "Login, authentication, permissions, or account settings",
"other": "Doesn't clearly fit any of the above",
},
),
"urgency": Score(
instructions="Rate how urgent this inquiry is.",
criteria=[
"low: a routine inquiry with no expectation of immediate response",
"medium: the customer is frustrated but there's no outage or hard deadline",
"high: outage, financial loss, strong demand for immediate action, or a hard deadline",
],
),
},
)
print(result.nouls["human_escalation"].noul)
print(result.choices["team"].choice)
print(result.scores["urgency"].score)
What a human designs here is the “shape of the question”:
- What to have the model decide
- What options to offer
- How to define each option
- What confidence threshold is safe to auto-act on
- Who to escalate to when confidence is low
The model, meanwhile, is the one actually deciding whether human_escalation is true, whether team is technical, and whether urgency is high.
Logging pays off later
Because Jev’s output is structured, it drops straight into your application logs as-is. Compare that output against human-confirmed ground truth later, and you get continuous, measurable accuracy tracking.
{
"request_id": "req_01H...",
"model": "jev-<pinned-version>",
"question_set_version": "support-routing-v3",
"state_hash": "sha256:...",
"answers": {
"human_escalation": {
"value": true,
"probability": 0.97
},
"team": {
"value": "technical",
"confidence": 0.81
},
"urgency": {
"value": 1.8,
"confidence": 0.76
}
},
"final_action": "handoff_to_human",
"human_label": {
"human_escalation": true,
"team": "technical",
"urgency": "high"
}
}
Here, urgency’s value is recorded on a 0-2 scale corresponding to low/medium/high.
At minimum, your logs should capture:
- The exact pinned model version used for inference
- The question-set version, including
instructionsandcriteria - The raw input text, or a hash / safely extracted features if you need to avoid storing PII
- Not just the selected value, but confidence or the full probability distribution
- The final action your code took
- The ground-truth label a human confirmed afterward
With logs and ground-truth labels in hand, you can track model performance and operational risk along several axes:
- Accuracy
- Precision, recall, and F1-score
- The confusion matrix, to spot which team or category gets misclassified most
- Auto-execution and mis-execution rates at each confidence threshold
- Calibration — checking whether a prediction at 0.9 confidence is actually right about 90% of the time
- Regression testing — confirming that a model or question-set update hasn’t degraded performance on the same eval set
Thresholds, calibration, and regression testing are beyond this article’s scope, so I’ll leave it at an overview. TypeSafe’s own recommendation is to route by confidence — auto-execute, confirm, or escalate to a human — and tune thresholds against your own data and the risk of your use case. For more, see TypeSafe: Confidence and Confidence-gated routing.
Worth noting: a confidence of 0.9 doesn’t guarantee that single prediction has a 90% chance of being correct. Whether 0.9 actually tracks real-world accuracy is something you can only determine by collecting enough examples for your specific use case and checking the actual accuracy rate within each confidence band.
The state of the official Playground
TypeSafe offers a Playground in its official Console, where you can type in text and try out Noul, Choice, and Score questions. The official Quick Start also points to the Playground as the entry point.
As of this writing, though, demand has spiked, and the sign-in screen showed this message in my environment:
Whoops, we're full
The URL included signups_disabled, confirming new account creation is paused.
Once I can get in, here’s what I plan to test:
- Classification accuracy on Japanese-language inquiries
- Output stability across repeated identical inputs
- How accuracy shifts with the granularity of
instructionsandcriteria - Latency when batching multiple questions into a single request
- Output differences between the real Jev and the local-compatible implementation
- How to design thresholds for routing low-confidence cases to human review
- Regression testing against a pinned model version
Trying it locally with jeff
The real Jev doesn’t publish its weights, so there’s no local execution or self-hosting option. You can’t fire up a real Jev with something like ollama run jev.
There is, however, a local-compatible implementation on GitHub that matches Jev’s API shape. The notable one is logan-markewich/jeff.
- GitHub: logan-markewich/jeff
- What it does: self-hosts a Jev-compatible
/v1/systemoneAPI - Backend: the open model GLiFormer (
knowledgator/gliformer-large-v1) - Notable: works with the official
typesafe-sdkas-is, so you can tryNoul,Choice, andScorelocally - Caveat: accuracy, confidence, and latency don’t match the real Jev
client = TypeSafeClient(
api_key="devkey",
base_url="http://localhost:8000",
)
Point base_url (or the TYPESAFE_BASE_URL environment variable) at your local server, and your application code can stay exactly the same while you work out question design, data shape, logging, and branch logic against a local instance.
jeff bills itself as a drop-in replacement, but the model running underneath it is a different one entirely. I’d treat it less as a substitute for Jev and more as a development and learning environment for trying out System One-style API design. If you want to evaluate the real model’s quality, confidence calibration, or operational SLA, you’ll need to wait for official API access and compare against the same eval dataset.
How to think about adoption
Right now, I think it’s too early to put Jev at the core of a production architecture. The reasons are straightforward:
- It just launched, with little long-term operational track record yet
- The Console’s signup pause is itself a signal worth watching — how well it scales with demand isn’t proven yet
- The real model is closed-weight; you can’t self-host it
- It’s not a fit as-is for systems handling sensitive data or PII you can’t send to an external API
- Its decisions are probabilistic — it doesn’t fully replace business rules or human review
- If you can’t fully trust your input data, you need to design for prompt injection and similar risks
That said, a PoC is worth doing if any of the following apply:
- You’re running high volumes of classification, routing, or relevance judgments over and over
- You’re repeatedly parsing and retrying a generative LLM’s JSON output
- You need low latency and confidence-based branching
- You can put together a small eval set with human-confirmed ground truth
If you do adopt it, abstract away the dependency on the official API so you can swap in a different implementation later. For example, put a layer like this in your application:
DecisionProvider
├─ TypeSafeJevProvider
├─ LocalJeffProvider
└─ LlmStructuredOutputProvider
That way, if Jev’s offering, pricing, accuracy, or usage limits change, you can switch implementations without rewriting your business logic.
Summary
TypeSafe AI’s Jev isn’t a text-generation LLM — it’s a System One Model built to return decisions inside your software as typed output.
What’s interesting isn’t just Jev’s performance as a single model. It’s the architecture behind it: instead of handing everything to a generative LLM, it splits the work by role.
Decision model (System One Model):
classification, routing, verification, scoring, guardrails
Generative LLM:
explanation, summarization, conversation, writing, complex reasoning
That split has real potential to improve cost, latency, reliability, and auditability when you’re putting AI workflows into production.
I’d start with a local-compatible implementation like jeff to work out question design, logging, an eval dataset, and threshold-based escalation. Once the real Jev’s Playground or API is available, compare accuracy, calibration, speed, and cost on the same data. Taken in that order, you can pull the System One design pattern into your own project without getting carried away by any one vendor’s hype.
References
TypeSafe official
- TypeSafe AI
- Introducing System One Models & Jev (TypeSafe AI Blog)
- TypeSafe Documentation
- Introduction: Jev and System One
- Quick Start
- How to build with TypeSafe
- System One
- Noul
- Choice
- Score
- Confidence
- API Reference
- TypeSafe Console / Playground
Local-compatible implementation
Evaluation metrics and operations
- Classification: Accuracy, recall, precision, and related metrics (Google Machine Learning Crash Course)
- Thresholds and the confusion matrix (Google Machine Learning Crash Course)
- Google Machine Learning Crash Course: Calibration
- scikit-learn: Probability calibration
- Google: Rules of Machine Learning