AI Atlas
All guides
⚡ GUIDE

System One Models: Jev, Clef and Laya — What They Do and Where to Use Them

The new class of models that make decisions in a single pass with calibrated probabilities instead of generating text: how they work, how Jev, Clef/Clef-flash and Laya differ, use cases, pairing them with an LLM, running locally, and pitfalls.

System One Jev Clef Laya Agents

Summary

A new class of models appeared in September–October 2026: System One models. They do not generate text. You give them a state (text, JSON, a ticket, an email; images for some) and typed questions, and the model returns the probability of every allowed answer in a single forward pass. Your code branches on those numbers.

This guide looks at the three notable examples: TypeSafe AI's Jev, which started the category; Cloudflare's open-weight Clef and Clef-flash; and Convai Innovations' small, locally run Laya. At the end you'll find a comparison table, a use-case catalog, an example of pairing with an LLM, and a checklist.

Note: Everything here was checked on October 5, 2026 against the vendors' own blog posts, documentation, Hugging Face model cards and GitHub READMEs. Almost all speed, cost and accuracy numbers are the vendors' own measurements; we say whose they are. The category is a few weeks old and details change fast.

What problem do they solve?

Software makes small decisions constantly: which team should get this ticket, is this message urgent, is this tool call safe, should this comment be moderated? Today we usually ask an LLM to "answer in this JSON format." That approach has three problems:

  • Slow and expensive: an LLM produces its answer token by token. Even a short JSON is dozens of sequential steps, plus invisible thinking tokens if reasoning is on.
  • Brittle: the output is text. You parse it, validate it, sometimes retry. Answers that break the schema or invent a non-existent option are possible.
  • Silent about uncertainty: you can ask an LLM "how sure are you?", but as TypeSafe's launch post puts it, models tend to be overconfident and inconsistent. A model that does a task right 95% of the time but can't tell you which 5% it gets wrong can't automate that task.

System One models target all three: you define the answer set up front, the model scores every option in one pass, and you branch on probability: "apply automatically", "ask a human" or "escalate to a bigger model".

How they work

Typed questions

All three models use the same three question types (TypeSafe's docs, Cloudflare's Workers AI schema and the Laya README define them the same way):

Type What it asks What it returns
choice Which of the options you defined? choice (most likely option), per-option probabilities, confidence
score Which level on an ordered rubric? score (probability-weighted level, can land between levels), per-level probabilities, confidence, legend
noul Is this yes/no statement true? noul: probability of "yes" (0–1)

One request carries several questions, all evaluated against the same state. The request body has the same skeleton across all three models:

{
  "model": "jev-latest",
  "state": "Hi, I've been trying to connect my Stripe account for 3 days and the integration keeps failing. I'm losing sales. Please help ASAP.",
  "questions": {
    "department": {
      "type": "choice",
      "instructions": "Which team should handle this",
      "criteria": {
        "billing": "Payment or subscription issues",
        "technical": "Bugs or integration problems",
        "sales": "Pricing or account questions"
      }
    },
    "frustration": {
      "type": "score",
      "instructions": "How frustrated the customer appears",
      "criteria": ["Calm, just stating facts", "Frustrated but civil", "Very angry, strong language"]
    },
    "is_urgent": { "type": "noul", "instructions": "The message conveys urgency or time-sensitivity" }
  }
}

In the sample response on TypeSafe's quick-start page, department comes back as "choice": "technical", "probabilities": {"technical": 0.85, "billing": 0.15, "sales": 0.0} and "confidence": 0.78. You pick the question keys (department, is_urgent); answers come back under the same keys.

Single pass

In an LLM, each new token is produced in sequence, conditioned on the previous ones. A System One model scores every option of every question at once. The architecture varies by vendor:

  • Clef adds a small "joint schema head" on top of a Qwen-based backbone; the head reads the backbone's final hidden states and emits one logit per option per question, turned into probabilities with a per-question softmax (Hugging Face model card).
  • Laya uses an encoder such as ModernBERT/mmBERT with a decision head on top (receptron/laya README).
  • Jev's architecture is not published; TypeSafe only says it uses "a new model architecture" and a "parallel sampler".

According to TypeSafe's docs, adding questions barely changes response time, and because each question is evaluated independently, more questions don't cause "context rot".

Calibration

A calibrated model should be right about 80% of the time when it says "0.8". TypeSafe says it uses a training method it calls RLCD (Reinforcement Learning for Calibrated Decisions). Laya uses the RLCD name too and says it is trained with reinforcement learning against strictly proper scoring rules. Cloudflare says Clef was trained with label-smoothed cross-entropy plus a Brier loss for calibration.

Those are vendor claims. Calibration depends on the distribution: don't trust it until you've measured it on your data, with your question wording, in your language. The checklist below shows how.

Confidence means different things per model

The API contract is shared, but the confidence formula is not. Jev (and TypeSafe's docs) use (n·p_max − 1)/(n − 1) for choice. Laya's README states that its confidence is "1 minus normalized entropy" and explicitly says thresholds carried over from Jev do not transfer. Re-measure thresholds whenever you switch models.

Jev (TypeSafe AI)

TypeSafe AI announced Jev in early access on September 15, 2026, in a launch post by founder Diogo Almeida. The model is closed-weight and only available through TypeSafe's API.

  • Endpoint: POST https://api.typesafe.ai/v1/systemone with Authorization: Bearer <API_KEY>. The model name is the jev-latest alias, which pointed to jev-1.13.0 as of October 5.
  • Price: $0.042 per million input tokens; output tokens are free (Models page).
  • Limits: 64k tokens per request; the state plus the longest question must fit in 32k. Text-only input (string, JSON object or array of text); no images, audio or video. Up to 255 options per choice. Rate limits are listed as 100K tokens/s and 80 requests/s, but TypeSafe says they may change without notice as demand grows.
  • Languages: English is the primary training language; other languages are "handled but not equally well". TypeSafe recommends testing on your own content before relying on Jev for non-English workloads.
  • Customization: Jev is not fine-tuned on customer data; every account uses the same weights. You adapt it to your domain through the state, instructions and criteria.
  • SDKs: Python (pip install typesafe-sdk) and JavaScript (@typesafe-ai/sdk) clients.

TypeSafe's claims: According to the launch post, Jev reaches similar intelligence to existing LLMs on System One tasks while being "two orders of magnitude" faster and more efficient, with 70–500 ms end-to-end response times. The post itself says the "193.6x faster, 444.6x cheaper" figures on the home page come from its own workflow evals and are "on the higher end of real world gains". TypeSafe also notes that it used the average of other companies' large models as the reference answer, and that its own team built the evals.

Known weaknesses: TypeSafe's "Jev 1.13 jaggedness" page is refreshingly candid: literal reading of instructions, counting and arithmetic, date comparison, multi-hop indirection, large states full of irrelevant detail, contradictory instructions/criteria, sensitivity to choice option order, and, of course, generation. Its advice is clear: do the math and date logic in code, and ask the model only for the judgment.

Clef and Clef-flash (Cloudflare)

On October 1, 2026, Cloudflare announced the first models trained by its Workers AI team: @cf/cloudflare/clef and @cf/cloudflare/clef-flash. Both belong to the same family as Jev and are Jev API-compatible; Cloudflare says you can move an existing Jev integration by changing the endpoint and model name.

  • Size and base model: Clef has 27B parameters, post-trained from Qwen3.8-27B; Clef-flash has 9B, from Qwen3.5-9B. According to the blog post, the backbone was frozen and rank-256 LoRA adapters were trained jointly with the decision head.
  • License: Open weights on Hugging Face under Apache 2.0 (Cloudflare/clef, Cloudflare/clef-flash).
  • Input: Text, JSON, images and video (model card). On Workers AI, images are sent embedded in an images field (max 4; remote URLs are not accepted).
  • Limits (Workers AI schema): 1–64 questions per request; 2–255 options per choice; 2–10 levels per score; 65,536-token context window.
  • Price (Workers AI): Clef $0.24 and Clef-flash $0.09 per million input tokens.
  • Fine-tuning: Cloudflare announced an RL fine-tuning service the same day; for now it runs as a design-partner program with Cloudflare's own engineers.

Cloudflare's claims: Across 43 benchmark runs, median latency was 209.3 ms for Clef, 38.8 ms for Clef-flash and 524.1 ms for Jev (p95: 238.6 / 122.4 / 536.0 ms). Cloudflare says a Clef model scores highest on 7 of 10 decision benchmarks, and that Clef beats Jev in 3 of 4 areas on TypeSafe's own workflow evals. The full table on the model card is more mixed: on reasoning-heavy tests such as GPQA Diamond, MMLU-Pro and BBH, Jev is clearly ahead. So it isn't "Clef wins everything"; it depends on the task.

Laya (Convai Innovations)

Laya is Convai Innovations' open-weight, Apache 2.0-licensed System One model. Its Hugging Face repo and the first laya release on PyPI are dated September 18, 2026. What sets it apart is size: hundreds of millions of parameters, not billions.

Checkpoint Encoder Params Context Best for
laya ModernBERT-large 421M 512 English
laya-multilingual mmBERT-base 322M 1,024 (up to 8,192) 100+ languages
laya-typed-decisions ModernBERT-large 421M 1,024 Typed-decision workflows (fine-tuned)
  • Router: the Router in the laya package detects the script and language and sends each request to the English or multilingual checkpoint.
  • Jev-compatible server: laya-serve exposes the same POST /v1/systemone contract as Jev; according to the README, an existing Jev client only needs its baseUrl changed.
  • Node/TypeScript: the @receptron/laya package runs the model through ONNX Runtime; no Python or PyTorch at runtime.
  • Fine-tuning: a fine-tuning notebook runs on Kaggle's free 2×T4 GPUs. According to the model card, the fine-tuned laya-typed-decisions lifts accuracy on the same 2,000 decisions from 0.362 (base English checkpoint) to 0.766.

Convai's claims: About 33 ms for one question and 7.2 ms per question batched, measured on a T4 GPU. According to the receptron README, a three-question call through the Node package takes about 140 ms on an Apple-silicon CPU once warm.

Honest limits (from the README itself): options share a fixed token budget (192 tokens on the English checkpoint, 256 on the others), so accuracy drops noticeably above roughly 20 options; on Banking77 (72–77 labels) Jev scores 0.870 and Laya 0.425. The English checkpoint is confidently wrong on non-Latin scripts, which is why the Router exists. No separate Turkish results are published; measure on your own data for Turkish workloads.

A note on independent comparison: in Cloudflare's own Decision Index run on the Clef model card, Laya trails Clef and Jev by a wide margin on most benchmarks, while having the lowest median latency at 5.8 ms. That run was done by a competitor and doesn't say which Laya checkpoint or settings were used; Convai's own tables look more favorable. Test both against your own evaluation.

Comparison table

Only cells verified from primary sources are filled; anything we couldn't verify is "—".

Jev Clef Clef-flash Laya
Vendor TypeSafe AI Cloudflare Cloudflare Convai Innovations
Released Sep 15, 2026 (early access) Oct 1, 2026 Oct 1, 2026 Sep 18, 2026 (first HF and PyPI release)
Size — (not published) 27B (Qwen3.8-27B base) 9B (Qwen3.5-9B base) 421M (English) / 322M (multilingual)
Weights / license Closed, API only Open, Apache 2.0 Open, Apache 2.0 Open, Apache 2.0
Deployment TypeSafe API Workers AI or your own GPU Workers AI or your own GPU Your own hardware: Python, laya-serve, ONNX/Node
Input Text/JSON only Text, JSON, images, video Text, JSON, images, video Text/JSON
Context 64k (state + longest question ≤ 32k) 65,536 tokens 65,536 tokens 512 / 1,024 (multilingual: up to 8,192)
Option limits 255 per choice 2–255; up to 64 questions per request 2–255; up to 64 questions per request ~20 options recommended; server cap 100
Languages English primary, others weaker — — 100+ (multilingual checkpoint)
Price $0.042 / M input tokens, output free $0.24 / M input tokens $0.09 / M input tokens Free (your hardware)
Latency claim TypeSafe: 70–500 ms end-to-end Cloudflare: 209.3 ms median Cloudflare: 38.8 ms median Convai: ~33 ms/question, 7.2 ms/question batched (T4)
Fine-tuning None (no per-customer weights) Cloudflare RL service (design partner) Cloudflare RL service (design partner) Open notebook (2×T4)

Use-case catalog

Use case Example questions Good fit Why
Support ticket triage team (choice), urgent (noul), frustration (score), churn_risk (noul) Laya (local, cheap), Clef-flash, Jev Few options, text-heavy, high volume; the headline scenario in all three vendors' docs
Agent tool gating / routing safe_to_run (noul), route (choice: code / search / human), needs_approval (noul) Clef-flash, Laya Runs on every step of the agent loop, so latency matters; Clef-flash's median is 38.8 ms, Laya runs locally in tens of ms
Content moderation policy (choice), severity (score), contains_threat (noul) Clef (if images are involved), Jev Clef and Clef-flash are the only option for image content; send low-confidence cases to a human
Fraud / risk flags suspicious (noul), risk_level (score) Clef, Jev Let the model judge, but keep amount and time arithmetic in code (a documented Jev weakness); confidence-gated human review is a must
Game NPC decisions action (choice: attack / hide / talk), threat (score) Jev, Clef-flash, Laya TypeSafe showed a Doom demo making 10 queries per second, with game state given as structured data rather than images
Robotics next_skill (choice), grasp_ok (noul) Clef / Clef-flash (vision), Laya (on-device) The Clef family is the only open option that reads camera frames directly; low-level control should stay in a classic controller
Form / document classification doc_type (choice), is_signed (noul), legible (noul) Clef (scanned documents), Jev / Laya (text) The Clef model card includes a receipt-image example asking "is the total legible?"
Lead scoring fit (score), budget_mentioned (noul), segment (choice) Jev, Laya Score with atomic questions and weight them in code (TypeSafe's "composite scoring" pattern)

Two rules of thumb: if you have more than about 20 options, look at Jev or Clef rather than Laya (or narrow the candidates first); if the input includes images or video, only the Clef family reads them directly today.

Pairing with an LLM: System 1 gates, System 2 reasons

The most efficient setup is usually this: a System One model looks at every request, settles most simple cases itself, and only sends what needs it to an expensive LLM or a human. Cloudflare positions Clef the same way: in the agent's request path, deciding first and handing off to an LLM to take action.

// Cloudflare Worker: Clef-flash gates, the LLM is only called when needed.
// handleRefund, callLLM and queueForHuman are your own functions.
const decision = await env.AI.run("@cf/cloudflare/clef-flash", {
  model: "clef-flash",
  state: { ticket: ticketText, plan: customer.plan },
  questions: {
    intent: {
      type: "choice",
      instructions: "What does the customer want?",
      criteria: {
        refund: "Asks for money back for a specific charge",
        how_to: "Asks how to use a feature",
        bug: "Reports something broken",
        other: "Anything else",
      },
    },
    abusive: { type: "noul", instructions: "Is the message abusive or threatening?" },
  },
});

const { intent, abusive } = decision.answers;

if (abusive.noul > 0.8) return queueForHuman(ticketText, "abuse");      // Safety: to a human
if (intent.confidence < 0.6) return callLLM(ticketText);                // Unsure: to System 2
if (intent.choice === "refund") return handleRefund(ticketText);       // Deterministic code
return callLLM(ticketText, { hint: intent.choice });                    // Writing replies is the LLM's job

The thresholds (0.8 and 0.6) are examples. Pick real values on a labelled validation set by trading off "share of cases handled automatically" against "error rate among automated cases". Sending uncertain cases to a person is a classic human-in-the-loop pattern; stopping harmful requests at the gate is a guardrail.

The same setup works in agent harnesses: a hook that asks "is this command destructive?" before every tool call, or a router that sends the user's message to the right sub-agent. TypeSafe's own "skill suggestion" and "guardrails for LLMs" cookbooks do exactly that.

Running locally

Laya: even on a laptop

pip install "laya[serve]"                     # Python 3.10+
LAYA_DEVICE=cuda LAYA_PRELOAD=1 laya-serve    # 0.0.0.0:8000, preloads all three checkpoints

curl -s localhost:8000/v1/systemone -H 'content-type: application/json' -d '{
  "state": {"body": "billed twice, refund please or we cancel"},
  "questions": {"dept": {"type": "choice", "instructions": "which team?",
                "criteria": {"billing": "refunds", "tech": "bugs"}}}
}'

Without a GPU, use LAYA_DEVICE=cpu (left unset, the device is picked automatically). The server binds to 0.0.0.0 by default; if it's reachable on a network, set LAYA_API_KEY so clients must send Authorization: Bearer <key>. According to the README, Router(preload=True) takes ~33 ms on GPU and 193–464 ms on CPU. On Node.js:

import { Laya } from "@receptron/laya"; // npm install @receptron/laya (Node 20+)

const laya = await Laya.load();
const result = await laya.systemOne(
  { subject: "Refund not received", body: "I cancelled two weeks ago and still have no refund..." },
  { churn_risk: { type: "noul", instructions: "Is the customer likely to cancel or dispute?" } },
);
console.log(result.answers.churn_risk.noul);
await laya.close();

Hardware note: the Hugging Face model card lists downloads of about 808 MB for the English checkpoint and about 647 MB for the multilingual one. The Node package downloads fp32 ONNX weights (~1.7 GB) and suggests budgeting ~2 GB of RAM. An ordinary server or laptop is enough.

Clef: you need a serious GPU

The reference code on Clef's model card was tested with torch 2.11 and transformers 5.10.2 on a single H200. Weights are bf16; the Hugging Face repo is about 55 GB for Clef and about 19 GB for Clef-flash. Our rough estimate (not measured): Clef needs at least an 80 GB data-center GPU, Clef-flash a 24 GB+ GPU; long states and batching need extra memory.

import sys
from huggingface_hub import snapshot_download

path = snapshot_download("Cloudflare/clef-flash")
sys.path.insert(0, path)  # the repo ships its own model code (joint_schema_model.py)
from joint_schema_model import load_release_model, systemone

model, processor = load_release_model(path, device="cuda")
response = systemone(model, processor, {
    "model": "clef-flash",
    "state": "Our checkout started returning errors and orders are blocked.",
    "questions": {
        "outage": {"type": "noul", "instructions": "Is a service down?"},
        "urgency": {"type": "score", "criteria": ["Can wait", "This week", "Today"]},
    },
})
print(response["answers"])

Because the repo ships its own Python code, read joint_schema_model.py before running it; this is remote code execution. Without your own GPU, the hosted version on Workers AI is far less hassle.

Limitations and pitfalls

  • Calibration drift: a model can be calibrated on its training distribution but not on your data, in your language, or a few months later when traffic shifts. Track reliability diagrams and Brier score regularly; Laya's own README documents its English checkpoint scoring 0% accuracy on Khmer at 0.952 confidence (the raw value before the temperature clamp; 0.705 as served).
  • Label design: the model answers what you asked, not what you meant. Option descriptions should be mutually exclusive and cover edge cases. Adding an "other" or "unclear" option keeps the model from being forced into a wrong category. TypeSafe also reports sensitivity to option order: shuffle the options and check the answer stays consistent.
  • Domain shift: Jev isn't fine-tuned per customer; Clef fine-tuning is still a partnership service; Laya you can fine-tune yourself. In very niche domains (legal or medical jargon, say) zero-shot results may disappoint.
  • Math, dates, counting: do them in code. Ask the model only for judgment.
  • Long, noisy states: irrelevant detail hurts accuracy. Filter the state per question; Laya's 512–1,024-token context is short anyway.
  • Benchmarks are the vendors': almost every number is the company's own run, and the companies measure each other with different settings. Compare on your own labelled set.
  • Maturity: the category is weeks old. Jev is in early access and its rate limits fluctuate; aliases (jev-latest) can silently move to a new version. Pin versions and log the model field.

Checklist

Fit

  • Can the answer set be defined up front (options, rubric or yes/no)?
  • Is the decision the kind an expert makes in a few seconds, or does it need multi-step reasoning?
  • Are arithmetic, counting and date comparisons in code?

Model choice

  • Can the data leave your infrastructure? If not: Laya, or Clef on your own GPU.
  • Any image/video input? If so: the Clef family.
  • More than ~20 options? If so: Jev or Clef, or narrow the candidates first.
  • Non-English content (e.g. Turkish)? Measure it separately on your own data.

Question design

  • Does each question ask one atomic thing?
  • Are options mutually exclusive, with an "other/unclear" option?
  • Is the answer stable when you shuffle option order?

Calibration and thresholds

  • Do you have a validation set of at least a few hundred labelled examples?
  • Were confidence thresholds tuned on this model with this wording (not carried over from another model)?
  • Do low-confidence cases go to a human or a larger LLM?

Production

  • Is the model version pinned, and is the response's model field logged?
  • Are calibration and the automated-decision rate monitored over time?
  • Is there retry and a fallback path for rate limits (429)?

Further reading