AI Atlas
EN TR
Intermediate · ~2 min read #decision-model #jev #clef

System One Model

Decisions in a single pass

A decision model that, instead of generating text, takes a state plus predefined typed questions and returns a calibrated probability for every allowed answer in a single forward pass.

SAME DECISION, TWO PATHSLLM · AUTOREGRESSIVE"Which team? Answer in JSON."{step 1"team"step 2:step 3"tech"step 4}step 5every token is another forward passJSON.parse() → validate → retry?one answer · no probabilitiesslow · needs parsingSYSTEM ONE · SINGLE PASSstate: "Checkout down for an hour"CHOICEteamNOULurgentSCOREseverity1 FORWARD PASS · ALL QUESTIONSteam = technical0.91team = billing0.07team = sales0.02urgent = yes0.96fast · typed · probabilisticThe LLM writes an answer; System One scores every allowed answer (values illustrative)
Definition

Ask a classic LLM "which team should handle this support ticket?" and it writes the answer token by token; you then parse and validate the text. A System One model never generates text. You give it a state (text, JSON, an email, a ticket; images for some models) and typed questions; it computes the probability of every allowed answer in one forward pass and returns them as structured values.

Today's models share three question types: choice (one of the options you define, plus a probability per option and a confidence), score (a probability-weighted level on an ordered rubric) and noul (the probability of "yes" for a yes/no question). Because the answer schema is fixed in advance, the output always matches those types; there is no such thing as a parse error.

The name comes from Daniel Kahneman's Thinking, Fast and Slow: System 1 is fast and intuitive, System 2 slow and deliberate. Reasoning models imitate System 2 by producing a long chain of thought before answering; System One models target narrow, well-scoped judgments of the kind an expert makes in a few seconds.

TypeSafe AI started the category with Jev in September 2026 (September 15, early access). Open-weight Laya (Convai Innovations) and Cloudflare's Clef / Clef-flash (October 1) followed; all three use the same state + questions → answers contract.

Analogy

Think of a triage nurse in an emergency room. She doesn't write an essay about each patient; in a few seconds she ticks three boxes: which department, urgency 1–5, "needs a doctor right now?" yes/no. And a good nurse knows when she is unsure and sends the case to a senior doctor.

Asking an LLM for a decision is like asking that nurse to write a paragraph for every patient and then extracting the boxes from the paragraph yourself. A System One model ticks the boxes directly and adds a "how sure am I" note to each one.

Real-world example

A message lands in a support queue: "Checkout has been failing for every customer for the last hour." One request asks three questions: team (choice: billing / technical / sales), urgent (noul) and severity (score: no impact → critical). The response contains the chosen team with a probability per team, the probability that it is urgent, and a probability-weighted severity score.

Your code branches on those numbers: if urgent > 0.9 and team = technical, page the on-call engineer; if confidence is low, hand the ticket to a human or a larger LLM. No free text is parsed anywhere.

Code examples
Clef · Cloudflare Workers AI typescript
// Adapted from the Cloudflare docs example (@cf/cloudflare/clef)
export default {
  async fetch(request, env): Promise<Response> {
    const response = await env.AI.run("@cf/cloudflare/clef", {
      model: "clef", // or "clef-flash" (@cf/cloudflare/clef-flash)
      state: "Checkout has been failing for every customer for the last hour.",
      questions: {
        urgent: { type: "noul", instructions: "Is this support request urgent?" },
        team: {
          type: "choice",
          instructions: "Which team should handle this request?",
          criteria: { billing: "Payments and refunds", technical: "Outages and errors", sales: "Plans and upgrades" },
        },
      },
    });
    // response.answers.urgent.noul    -> probability of "yes" (0–1)
    // response.answers.team.choice    -> most likely team (+ probabilities, confidence)
    return Response.json(response);
  },
} satisfies ExportedHandler<Env>;
Jev · TypeSafe HTTP API bash
curl -X POST https://api.typesafe.ai/v1/systemone \
  -H "Authorization: Bearer $TYPESAFE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "jev-latest",
    "state": "Help! My payouts have been failing for 3 days.",
    "questions": {
      "is_urgent": { "type": "noul", "instructions": "Does this convey urgency?" }
    }
  }'
# Response: {"model":"jev-1.13.0","answers":{"is_urgent":{"type":"noul","noul":...}},
#            "usage":{"input_tokens":...,"output_tokens":...}}
Laya · local, open weights python
# pip install laya
from laya import Router

router = Router()  # downloads a checkpoint from Hugging Face on first use
state = "Hi, we were billed twice for March. Please refund the duplicate today or we will cancel our plan."
questions = {
    "department": {"type": "choice", "instructions": "Which department should handle this?",
                   "criteria": {"billing": "payments, refunds", "technical": "bugs, outages", "other": "everything else"}},
    "churn_risk": {"type": "noul", "instructions": "Does the user threaten to cancel or leave?"},
}

result = router.predict(state, questions)
print(result["answers"]["department"]["choice"])  # e.g. billing
print(result["answers"]["churn_risk"]["noul"])    # probability of "yes"
print(result["routing"]["model"])                 # chosen checkpoint: english / multilingual
When to use
  • Decisions with a known answer set: routing, triage, intent classification, risk flags
  • Fast gates inside an agent loop: is this tool call safe, which sub-agent should take it, should a human be asked?
  • Latency-critical hot paths: game NPCs, robot control loops, per-request moderation
  • Anywhere you want to branch on probability: automate the confident cases, escalate the uncertain ones to a human or a bigger model
  • Cheaply labelling, scoring or extracting features from large numbers of records
When not to use
  • Open-ended generation: writing replies, summarizing, translating, writing code — these models do not generate text
  • Multi-step reasoning, math, date comparison, counting — TypeSafe's own docs say to keep these in code
  • Questions whose answer set cannot be defined up front
  • Narrow tasks where you already have labelled data and a classic classifier that is good, cheap and auditable enough
Common pitfalls

Mistaking probability for accuracy

These models are trained for calibration, but calibration cannot be trusted until you measure it on your own data and question design. Laya's own README documents its English checkpoint being confidently wrong on non-Latin scripts.

Carrying confidence thresholds across models

Same API, different confidence: Jev uses (n·p_max − 1)/(n − 1) while Laya uses 1 − normalized entropy. A threshold tuned on Jev does not transfer to Laya or any other model; re-measure per model.

Cramming everything into one question

System One models do well on narrow, atomic questions. Instead of 'is this startup good?', ask about market, technical feasibility and differentiation separately, and do the weighting in code.

Sloppy option labels

The model reads your words literally. Similar or overlapping option descriptions and very many options hurt accuracy; on Laya, options share a fixed token budget and a clear drop above roughly 20 options is documented.

Forgetting the category is brand new

Jev is in early access; Clef and Laya are weeks old. Most benchmarks are the vendors' own. Pin a model version (e.g. jev-1.13.0) instead of an alias and decide with your own evaluation set.