System One Model
Decisions in a single pass
A decision model that, instead of generating text, takes a state plus predefined typed questions and returns a calibrated probability for every allowed answer in a single forward pass.
Ask a classic LLM "which team should handle this support ticket?" and it writes the answer token by token; you then parse and validate the text. A System One model never generates text. You give it a state (text, JSON, an email, a ticket; images for some models) and typed questions; it computes the probability of every allowed answer in one forward pass and returns them as structured values.
Today's models share three question types: choice (one of the options you define, plus a probability per option and a confidence), score (a probability-weighted level on an ordered rubric) and noul (the probability of "yes" for a yes/no question). Because the answer schema is fixed in advance, the output always matches those types; there is no such thing as a parse error.
The name comes from Daniel Kahneman's Thinking, Fast and Slow: System 1 is fast and intuitive, System 2 slow and deliberate. Reasoning models imitate System 2 by producing a long chain of thought before answering; System One models target narrow, well-scoped judgments of the kind an expert makes in a few seconds.
TypeSafe AI started the category with Jev in September 2026
(September 15, early access). Open-weight Laya (Convai Innovations)
and Cloudflare's Clef / Clef-flash (October 1) followed; all three
use the same state + questions → answers contract.
Think of a triage nurse in an emergency room. She doesn't write an essay about each patient; in a few seconds she ticks three boxes: which department, urgency 1–5, "needs a doctor right now?" yes/no. And a good nurse knows when she is unsure and sends the case to a senior doctor.
Asking an LLM for a decision is like asking that nurse to write a paragraph for every patient and then extracting the boxes from the paragraph yourself. A System One model ticks the boxes directly and adds a "how sure am I" note to each one.
A message lands in a support queue: "Checkout has been failing for
every customer for the last hour." One request asks three questions:
team (choice: billing / technical / sales), urgent (noul) and
severity (score: no impact → critical). The response contains the
chosen team with a probability per team, the probability that it is
urgent, and a probability-weighted severity score.
Your code branches on those numbers: if urgent > 0.9 and team =
technical, page the on-call engineer; if confidence is low, hand the
ticket to a human or a larger LLM. No free text is parsed anywhere.
// Adapted from the Cloudflare docs example (@cf/cloudflare/clef)
export default {
async fetch(request, env): Promise<Response> {
const response = await env.AI.run("@cf/cloudflare/clef", {
model: "clef", // or "clef-flash" (@cf/cloudflare/clef-flash)
state: "Checkout has been failing for every customer for the last hour.",
questions: {
urgent: { type: "noul", instructions: "Is this support request urgent?" },
team: {
type: "choice",
instructions: "Which team should handle this request?",
criteria: { billing: "Payments and refunds", technical: "Outages and errors", sales: "Plans and upgrades" },
},
},
});
// response.answers.urgent.noul -> probability of "yes" (0–1)
// response.answers.team.choice -> most likely team (+ probabilities, confidence)
return Response.json(response);
},
} satisfies ExportedHandler<Env>;curl -X POST https://api.typesafe.ai/v1/systemone \
-H "Authorization: Bearer $TYPESAFE_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "jev-latest",
"state": "Help! My payouts have been failing for 3 days.",
"questions": {
"is_urgent": { "type": "noul", "instructions": "Does this convey urgency?" }
}
}'
# Response: {"model":"jev-1.13.0","answers":{"is_urgent":{"type":"noul","noul":...}},
# "usage":{"input_tokens":...,"output_tokens":...}}# pip install laya
from laya import Router
router = Router() # downloads a checkpoint from Hugging Face on first use
state = "Hi, we were billed twice for March. Please refund the duplicate today or we will cancel our plan."
questions = {
"department": {"type": "choice", "instructions": "Which department should handle this?",
"criteria": {"billing": "payments, refunds", "technical": "bugs, outages", "other": "everything else"}},
"churn_risk": {"type": "noul", "instructions": "Does the user threaten to cancel or leave?"},
}
result = router.predict(state, questions)
print(result["answers"]["department"]["choice"]) # e.g. billing
print(result["answers"]["churn_risk"]["noul"]) # probability of "yes"
print(result["routing"]["model"]) # chosen checkpoint: english / multilingual- Decisions with a known answer set: routing, triage, intent classification, risk flags
- Fast gates inside an agent loop: is this tool call safe, which sub-agent should take it, should a human be asked?
- Latency-critical hot paths: game NPCs, robot control loops, per-request moderation
- Anywhere you want to branch on probability: automate the confident cases, escalate the uncertain ones to a human or a bigger model
- Cheaply labelling, scoring or extracting features from large numbers of records
- Open-ended generation: writing replies, summarizing, translating, writing code — these models do not generate text
- Multi-step reasoning, math, date comparison, counting — TypeSafe's own docs say to keep these in code
- Questions whose answer set cannot be defined up front
- Narrow tasks where you already have labelled data and a classic classifier that is good, cheap and auditable enough
Mistaking probability for accuracy
These models are trained for calibration, but calibration cannot be trusted until you measure it on your own data and question design. Laya's own README documents its English checkpoint being confidently wrong on non-Latin scripts.
Carrying confidence thresholds across models
Same API, different confidence: Jev uses (n·p_max − 1)/(n − 1) while Laya uses 1 − normalized entropy. A threshold tuned on Jev does not transfer to Laya or any other model; re-measure per model.
Cramming everything into one question
System One models do well on narrow, atomic questions. Instead of 'is this startup good?', ask about market, technical feasibility and differentiation separately, and do the weighting in code.
Sloppy option labels
The model reads your words literally. Similar or overlapping option descriptions and very many options hurt accuracy; on Laya, options share a fixed token budget and a clear drop above roughly 20 options is documented.
Forgetting the category is brand new
Jev is in early access; Clef and Laya are weeks old. Most benchmarks are the vendors' own. Pin a model version (e.g. jev-1.13.0) instead of an alias and decide with your own evaluation set.