The Question Nobody Asks

Somewhere in your agent's loop, right now, there's a call that looks like this: hand the full model a tool call, ask it to reason in prose about whether the call is safe, then parse a decision back out of whatever it wrote. Multi-second latency. A full generation bill. All to answer a question with exactly two possible answers.

Nobody designed this on purpose. It's just what happens when "ask the LLM" is the only tool in the drawer. Routing, risk-scoring, best-of-N selection, LLM-as-judge — a huge share of what agents spend tokens on isn't generation at all. It's classification wearing a chat costume.

A new class of model exists specifically to strip the costume off: it returns a typed decision, with a calibrated confidence score, and never generates a single token to get there.

Two entrants ship it right now, built on completely different mechanisms. One is closed and proprietary. The other is open-weight and came out eight days later, explicitly benchmarked against the first. Together they're the clearest evidence yet that "bigger model" isn't the only lever left to pull.


Two Ways to Skip the Paragraph

The first mover calls itself a "System One" model — Kahneman's fast, intuitive System 1 versus the deliberate, generative System 2 a chat LLM runs. You declare a typed schema up front (which fields, which allowed values per field), hand it program state, and it returns those fields filled in with calibrated probabilities attached, computed in one parallel pass. No decoding loop, no next-token sampling. Its training method is proprietary — the company hasn't published the architecture, weights, or training recipe.

The second entrant takes the opposite bet on openness and a completely different mechanism. Instead of any form of decoding, it embeds the current state and every candidate action into a shared vector space and picks whichever action's embedding sits closest to the state's — literally nearest-neighbor search, not generation of any kind. It's built as a dual-encoder contrastive model, architecturally closer to CLIP than to any chat LLM: a frozen base model plus two small trainable projection heads (roughly 20M parameters each), one for state, one for actions. Full weights, training data composition, and code are public under an open license, self-hostable on a single GPU.

Call the first one Jev and the second CLM. Where an example below is illustrating the pattern generically rather than one implementation specifically, call it a decision model — the category name for both.

The mechanism difference matters more than it sounds. Jev's calibration is trained directly — a reinforcement-learning method tuned so that "80% confident" actually means correct 80% of the time. CLM's confidence is a similarity score from contrastive training (InfoNCE, the loss family behind CLIP): it tells you which candidate is most like a correct answer relative to the others in the batch, which is a related but distinct property from calibration. That's the tradeoff running underneath everything below.


Under the Hood

The high-level pitch is the same for both: typed decision instead of generated text. The mechanisms underneath it aren't, and the gap explains almost everything that follows.

Jev runs non-autoregressive, parallel sampling: it produces every field of the declared output schema at once, not token by token, which is where its latency floor comes from. Its maker claims that because the output space is a fixed schema rather than free text, a type error is mathematically impossible — there's no unstructured generation step where malformed output could sneak in. Training is the proprietary RL method mentioned above, tuned specifically for calibration rather than for sounding plausible. End-to-end latency sits in the 70–500ms range, at a price close to two orders of magnitude below a flagship chat model per call. None of that is independently checkable, though — no architecture diagram, parameter count, or paper is public. Every number in this paragraph is the vendor's own disclosure.

CLM works nothing like it. It's a dual-encoder contrastive model, closer in spirit to CLIP than to any chat LLM:

The base model's weights stay frozen; only two small projection heads (roughly 20 million parameters each, a few dozen megabytes) get trained — one for state, one for actions. The frozen base already understands language; the heads just learn to place related state/action pairs near each other in vector space, which is what keeps training cheap. The objective is bidirectional InfoNCE, the same contrastive-loss family CLIP uses: shown one correct pair plus several wrong ones per step, rewarded for scoring the correct pair highest.

Training ran in three stages, and accuracy climbs visibly at each one: roughly 60 million question-answer pairs got it to 52% top-1 accuracy on held-out questions; adding about 30 million synthetic hard negatives pushed that to 69%; a final pass over roughly a million real agent trajectories turned it from a general-purpose matcher into something that behaves like an actual agent component. The whole thing self-hosts on a single GPU:

client = DecisionModelClient()
result = client.classify(state="...", schema={"is_safe": bool})

Caching is the design's whole point, not an afterthought: because state and action embeddings are computed independently, a fixed action set's embeddings get computed once and reused for every future decision. Per-call cost becomes "one fresh embedding plus a cached dot product per candidate" — exactly why the speed advantage grows, rather than shrinks, as the action space gets larger.


What a Typed Decision Actually Looks Like

The output shape is the whole point. Compare what a caller has to do with each approach:

response = llm.chat(
    f"Given this tool call: {tool_call}\n"
    "Should this be blocked? Reply with JSON: "
    '{"block": true/false, "confidence": 0-1, "reason": "..."}'
)

# Now hope it's valid JSON, hope the fields are named
# right, hope it didn't wrap it in markdown fences.
decision = json.loads(extract_json(response.text))

if decision["block"]:
    raise ToolBlocked(decision["reason"])

The "before" version isn't hypothetical — it's the default shape of tool-call gating in most agent codebases today. The "after" version doesn't just save tokens. It removes an entire class of bug: the response that's almost valid JSON.


The Numbers

Both models were benchmarked zero-shot against each other across gaming, tool-calling, and long-horizon navigation tasks. The latency gap is the headline, and it's consistent across every task:

Decision latency per call, ms — lower is better
Dinosaur Run — CLM16.5ms
Dinosaur Run — Jev149.8ms
Super Mario — CLM33.5ms
Super Mario — Jev132.6ms
WikiRacing — CLM79.8ms
WikiRacing — Jev225ms
Tool calling (BFCL v4) — CLM76.8ms
Tool calling (BFCL v4) — Jev125.5ms

The decision model runs 4–9x faster on every single task, and on the two pure-gaming tasks it matches Jev's success rate exactly (5/5 on both). The gap narrows where it costs the decision model accuracy: WikiRacing (26 of 30 tasks completed, versus a clean 30 of 30), and tool calling (95.2% versus 99.2%). Zero-shot, the decision model is trading a small amount of out-of-the-box accuracy for a large latency win.

That trade flips once CLM gets fine-tuned for a specific job. Used as a verifier — reranking candidate outputs from a coding agent rather than picking a game move — it overtakes Jev outright:

CLM (fine-tuned)
81.6% success
Jev
71.1% success

Same pattern shows up on a terminal-agent benchmark: 87.6% for the fine-tuned decision model against 83.1% for Jev, at 4–6x lower latency per decision either way. Worth saying plainly: these are the two labs' own released evaluations, run zero-shot and after fine-tuning respectively — not yet independently reproduced by a third party. Treat the direction of the result as more solid than the exact decimal.


A Worked Example: One Session's Token Bill

Percentages are convincing right up until someone asks what that means for an actual session. Here's the arithmetic for a plausible case: a coding agent working a multi-step task that needs a safety gate check before every tool call — the same block/allow shape from the code example above. A session like that typically clears somewhere around 40 tool calls before it's done.

Gating every call by asking the full chat model (the "before" path):

Per check × 40 calls
Input tokens ~550 (tool call + surrounding state + reasoning instructions) 22,000
Output tokens ~120 (a sentence or two of reasoning, then the JSON) 4,800
Latency ~2–3s — a real generation call 80–120s of the session spent purely on gating
Cost, at representative flagship pricing — roughly $0.14, before a single line of code gets written

Gating every call with a decision model (the "after" path):

Per check × 40 calls
Input tokens ~180 (state + typed schema, no reasoning-instruction bloat) 7,200
Output tokens 0 — a typed field and a probability, not prose 0
Latency 50–150ms 2–6s total
Cost, at the pricing profile covered above — a fraction of a cent

Same 40 decisions, same underlying question at each one. Raw token volume drops by roughly 70% before price even enters the picture — a "yes/no plus a number" answer was never going to need 120 tokens of prose to say it — and the tokens it does use cost close to two orders of magnitude less per unit. Stack that across every session an agent runs in a day, and a gate-checking layer goes from a real, visible line item to a rounding error. That's the whole argument compressed into one session's worth of arithmetic: it isn't just faster per call, it's cheaper by two independent multipliers stacked on top of each other — fewer tokens, and each token costs less.


Where This Actually Plugs Into an Agent

Four places this pattern is already showing up in production agent loops.

1. Inside a single decision — before vs. after

Same tool call, same question ("should this be allowed?"), two different loops for answering it — this is the code example from earlier, redrawn as a diagram:

flowchart LR subgraph Naive["🐌 Naive — the LLM answers every check itself"] U1[Tool call arrives] --> LLM1["LLM generates prose:
reasons, decides, explains"] LLM1 --> D1["Parse the decision back
out of that prose"] D1 --> Act1[Allow / block the action] end subgraph TwoSpeed["⚡ Two-speed — a classifier answers, the LLM only reasons once"] U2[Tool call arrives] --> LLM2["LLM reasons once, proposes
state + candidate actions"] LLM2 -->|state + candidates| SO["Decision model scores them:
typed value + confidence"] SO -->|typed value, nothing to parse| Act2[Allow / block the action] SO -.->|confidence too low, escalate| LLM2 end classDef navy fill:#bed4ef,stroke:#4285d7,color:#10161c classDef teal fill:#b1eaf1,stroke:#00b2ca,color:#10161c classDef mint fill:#c6ece0,stroke:#7dcfb6,color:#10161c class LLM1,LLM2 navy class D1,SO teal class Act1,Act2 mint

Blue is where an LLM does the work, teal is where the decision gets made, green is the action that decision produces. The naive path burns a full generation on every single check (~2–3s, ~670 tokens each); the two-speed path spends that generation once to propose candidates, then reuses the cheap classifier for each check after (~50–150ms, ~180 tokens each) — falling back to the LLM (the dashed edge) only when the classifier isn't confident enough to trust on its own.

The low-confidence escalation edge isn't a documented feature of either model — it's the natural consequence of a calibrated confidence score existing at all. If the score is honest, low confidence is exactly the signal a caller should use to fall back to the slower, smarter path instead of trusting the fast one blindly.

2. Inside an agent graph — routing and tool-call gating

flowchart TD Start([Incoming request]) --> Router{{Decision model: classify + route}} Router -->|routine, high confidence| Small[Cheaper model node] Router -->|ambiguous, high stakes| Big[Flagship LLM node] Small --> Plan[Agent proposes a tool call] Big --> Plan Plan --> Gate{{Decision model: score call risk}} Gate -->|low risk, authorized| Tool[Execute tool] Gate -->|risky or under-authorized| Block[[Blocked — no execution]] Tool --> More{More steps needed?} More -->|yes| Plan More -->|no| End([Return result]) classDef navy fill:#bed4ef,stroke:#4285d7,color:#10161c classDef amber fill:#fbd1a2,stroke:#f49a34,color:#10161c classDef orange fill:#fbd0b6,stroke:#f79256,color:#10161c class Router,Gate navy class Small,Big,Plan,Tool amber class Block orange

Both the routing decision and the gating decision get logged with their full classification and confidence in agent state — a trace shows exactly why a model was picked or a call was blocked, not just that it happened.

3. Best-of-N candidate selection

flowchart LR State[Current state / task] --> Gen[LLM generates N candidates] Gen --> C1[Candidate 1] Gen --> C2[Candidate 2] Gen --> C3[Candidate N ...] Cache[(Cached candidate embeddings)] -.->|reused across calls| Select C1 --> Select{{Decision model: score vs. state}} C2 --> Select C3 --> Select Select -->|highest similarity| Best[Best candidate selected] Best --> Exec[Agent applies it] classDef navy fill:#bed4ef,stroke:#4285d7,color:#10161c classDef teal fill:#b1eaf1,stroke:#00b2ca,color:#10161c classDef mint fill:#c6ece0,stroke:#7dcfb6,color:#10161c class Gen navy class C1,C2,C3,Cache teal class Select,Best,Exec mint

This is the shape behind the strongest fine-tuned numbers above: the LLM still does the actual generation — the decision model never invents a candidate — it only ever ranks what the LLM already proposed. Caching pays off hardest here, since the same candidate set (or overlapping candidates across similar tasks) gets re-scored repeatedly, and cached action embeddings don't need recomputing.

One independent test of exactly this caching behavior pushed it further: scoring roughly a thousand candidates at once, a bi-encoder design like this ran about 13x faster than a cross-encoder alternative that has to score every candidate jointly rather than independently. That's the concrete mechanism behind "the advantage grows with the size of the action space" — not a theoretical claim, a measured one.

4. LLM-as-judge

flowchart LR Agent[Agent under test] --> Trace[Captured trace / output] Trace --> JudgeOld[Full LLM judge — multi-second, higher cost] Trace --> JudgeNew[Decision model judge — sub-second, cheap] JudgeOld --> ScoreOld[Score — wobbles across repeated runs] JudgeNew --> ScoreNew[Score — stable across repeated runs] ScoreNew --> CI[Cheap enough to run inline in CI] classDef navy fill:#bed4ef,stroke:#4285d7,color:#10161c classDef orange fill:#fbd0b6,stroke:#f79256,color:#10161c classDef mint fill:#c6ece0,stroke:#7dcfb6,color:#10161c class JudgeOld,ScoreOld orange class JudgeNew,ScoreNew,CI mint class Agent,Trace navy

This is the use case with the loudest measured result: teams that swapped a chat-LLM judge for Jev specifically have reported 90x+ lower score variance on identical repeated evaluations, at a small fraction of the per-run cost — the difference between "safe to run inline in CI" and "overnight batch job only." The pattern generalizes to any decision model sitting in the judge seat; the number above is Jev's own reported result, not yet matched by a published CLM figure.


CLM vs Jev, Side by Side

Dimension Jev-style CLM-style
Decision mechanism Non-autoregressive parallel sampling over a typed schema Contrastive embedding similarity — nearest-neighbor over cached vectors
Training method Proprietary RL, tuned for calibration — unpublished Bidirectional InfoNCE contrastive loss — published, standard
Openness Closed — no public weights, paper, or architecture diagram Open — self-hostable, full weights and code available
Base model Undisclosed Frozen open base model + small trainable projection heads
Self-hostable? No — API access only, waitlist gated Yes — a single GPU is enough
Pricing Roughly two orders of magnitude below flagship chat pricing per call Free to self-host; you pay for the GPU, not per call
Generates new actions? No — typed schema only No — must select from a supplied candidate set
Caches action representations? Not part of the public design Yes — the core design principle
Calibration Trained directly toward calibrated confidence Relative similarity score, not inherently calibrated
Best zero-shot fit General-purpose typed decisions, no fine-tuning needed Large or repeatedly-visited action spaces where caching compounds
Best fine-tuned fit Still solid, doesn't need task-specific tuning to be strong Verifier/reranker roles — this is where it overtakes the closed model
Independent verification None — every number above is vendor-disclosed Partial — third parties reproduced the speed claims; the strongest fine-tuned accuracy numbers are still self-reported

Already Showing Up in Production

This isn't purely theoretical. Both sides of this comparison have real integration work behind them, within days of launch.

On the closed side: a popular open-source agent-orchestration framework shipped middleware built directly around this pattern almost immediately. One component classifies a request once per turn to route it to the right model, storing the full classification and confidence in agent state so the routing decision stays traceable after the fact. A second classifies each tool call for risk before it executes and returns a hard error instead of letting anything judged risky or under-authorized run — a pre-execution gate, not an after-the-fact flag. The same idea also shows up as an LLM-as-judge alternative inside at least one evaluation platform: swapping a chat-model judge for a typed-decision judge, on the same benchmark and the same dataset, cut score variance by 90–900x+ and dropped the cost of a full evaluation run from roughly $28 to about $0.34.

The dollar case scales further at the enterprise end. One study modeling a 10,000-seat organization running coding agents found that inserting a cheap classifier in front of the existing model-routing logic — deciding "does this request actually need the flagship model" before it reaches the vendor's own defaults — recovered 14–21% of total model spend. At that scale, that's several million dollars a year, from a component that never generates a single token.

On the open side, adoption is earlier but concrete rather than speculative. The DeepSWE and terminal-agent reranking numbers from earlier in this post aren't benchmark curiosities — that's the actual "LLM generates, CLM selects" pattern already running end to end. At least one existing decision-scoring toolkit in this space has opened its own integration effort, proposing to combine CLM's speed with a calibration objective it doesn't natively have (more on that gap below). A larger, multimodal successor is already reported to be in training.


A Landscape Forming Fast

CLM isn't the only attempt at an open answer to the closed model. At least seven other open-source projects are independently working the same angle — non-autoregressive, typed output instead of chat:

Project Architecture Size Key distinction
Laya Encoder + decision head 322M–421M Multilingual — 100+ languages
Nimble Small LLM + LoRA 9B Fully open: model, training recipe, eval harness, and data pipeline
Kev LLM + LoRA + pointer head 0.8B–9B Implements the closed model's own API contract directly — a drop-in swap
SemIf Direct logit extraction from open LLMs Model-dependent Reads the logits already assigned to permitted answers — no extra training needed
Rizzo Flow Small LLM + local serving 1.7B–4B Mirrors the closed model's HTTP interface
Von Encoder + custom scoring head ~395M Runs on CUDA, ROCm, Apple Silicon, or plain CPU
NanoJev Small LLM + shared decision heads 0.6B The smallest of the group — a sub-1B replica

None of these seven have published benchmark numbers against either model covered above. CLM is currently the most credible open attempt of the group, mostly because it's the only one backed by a named, methodical benchmark against the closed model rather than an unverified README claim.


The Trade You're Actually Making

💡 The ceiling to know about
Neither model generates a new action. It only ranks or classifies against a fixed, supplied set of options. If every candidate handed to it is wrong, it has no way to notice that and produce a correct one instead — that's still the LLM's job. This is a structural boundary, not a training gap, and it's the reason the right architecture is generation and decision split into two stages, never decision-only end to end.

There's a second, sharper distinction worth naming plainly: only one of these two was actually trained to be calibrated. Jev's training method targets calibration directly — a stated confidence is supposed to track real accuracy. CLM's training objective has no calibration term at all; it optimizes relative similarity within a batch, which is related but not the same property. A softmax score that looks like a probability isn't automatically one. If a pipeline is going to make automated calls off a confidence threshold, that distinction is worth checking, not assuming.

A third trade: a calibrated-sounding confidence score is still just a number a system trusts by default, and any model reading untrusted content to produce that number can, in principle, have its answer shifted by content written specifically to steer it. One published test made this concrete: asked whether to block a destructive shell command, the closed model returned roughly a 76% probability of blocking it — until a fake pre-approval was slipped into the tool output ahead of the check, after which the block probability fell to about 48%. The stated confidence looked reasonable in isolation; it wasn't robust to adversarial content in what it was reading. That's the same class of risk LLM-as-judge setups already carry, just with a smaller, faster model in the judge's seat instead of a bigger one.

Treat either model's verdict the way you'd treat any single automated check: useful, fast, and never the only gate standing between an agent and something destructive. Neither vendor's own numbers have been independently reproduced yet, either — the closed model's speed and cost claims rest entirely on its maker's own disclosures, and the open model's strongest fine-tuned results were measured on held-out evaluation subsets rather than full public benchmark submissions. Trust the direction of both stories. Treat the exact decimals as a starting point for your own measurement, not a finished one.


The Actual Idea

The headline isn't "a faster model." It's a question: does an agent need to generate text every time it has to make a choice? When the valid action set is already known — a fixed tool list, a fixed set of candidate patches, a fixed set of navigation links — the answer can be no. Vector-space selection, or a typed classifier, is enough, and it's an order of magnitude cheaper.

That reframes the agent loop as a two-speed system: an LLM that generates and reasons when reasoning is actually required, and a small, fast, purpose-built model that picks — with a real number attached to how sure it is — everywhere else. Because the second model's outputs can be cached, the pattern gets more valuable exactly where you'd expect: agents operating over a large, stable, or repeatedly-visited set of tools, patches, or links, where the caching advantage compounds every time it's reused instead of amortizing away.


Conclusion

The pattern here isn't "use this specific model" — it's a question worth asking at every decision point in an agent: is this actually a generation problem, or is it a bounded-choice problem wearing a generation costume?

Key takeaways:

  • A large share of what "agentic" LLM calls do isn't generation at all — it's classification with extra prose wrapped around it.
  • Two model classes now exist purpose-built to skip that prose entirely: one closed and general-purpose, one open and built around cached contrastive embeddings.
  • Zero-shot, the open approach runs 4–9x faster for a small accuracy cost; fine-tuned as a verifier, it can overtake the closed one on both speed and accuracy.
  • Neither replaces the reasoning model — both are built to sit next to one, absorbing the routing, gating, reranking, and judging calls that never needed a paragraph in the first place.