Skills DirectorySkills Directory
SkillsLearnSecurityCategoriesDocsBlogPro
Sign InSubmit Skill
Skills Directory

Security-tested agent skills for Claude, coding agents, and AI workflows.

Directory

  • Browse Skills
  • All Skills A–Z
  • Claude Skills
  • Claude Code Skills
  • Agent Skills
  • Categories
  • Authors
  • Submit a Skill

Learn

  • Learn Hub
  • Install Claude Skills
  • Write SKILL.md
  • Skills vs MCP
  • Directories Compared

Security

  • Security
  • Methodology
  • Secure Claude Skills
  • Security Badges
  • Chrome Extension
  • Skill Manager

Company

  • About
  • Community
  • Blog
  • API Docs
  • Advertise

2026 Skills Directory. All rights reserved.

ProTermsPrivacyRefunds
Back to skills

Typed Decision Models

ASecurity

Replacing an LLM used as a classifier with a model that returns typed decisions and probabilities instead of prose — TypeSafe's Jev and the wider pattern. Covers the three primitives, the decomposition that makes them work, what `confidence` does and does not mean, and the independently measured evidence on where the vendor's claims hold. USE WHEN: user mentions "Jev", "TypeSafe AI", "System One model", "typed decision model", "decision model", "calibrated probabilities", "classifier instead...

40 stars
0 votes
0 copies
0 views
Added 10/3/2026
ai-agentspythonrustrailsawsgitapibackend

Works with

cliapi

Security Analysis

A100/100

Scanned 10/3/2026

$npx -y skills add claude-dev-suite/claude-dev-suite --skill typed-decision-models --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Typed Decision Models?

Add the live security badge to your README — it updates automatically with every re-scan.

Security grade badge for Typed Decision Models
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/claude-dev-suite-typed-decision-models/badge)](https://www.skillsdirectory.com/skills/claude-dev-suite-typed-decision-models)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
Files
SKILL.md
---
name: typed-decision-models
description: |
  Replacing an LLM used as a classifier with a model that returns typed
  decisions and probabilities instead of prose — TypeSafe's Jev and the wider
  pattern. Covers the three primitives, the decomposition that makes them work,
  what `confidence` does and does not mean, and the independently measured
  evidence on where the vendor's claims hold.

  USE WHEN: user mentions "Jev", "TypeSafe AI", "System One model", "typed
  decision model", "decision model", "calibrated probabilities", "classifier
  instead of LLM", "logprobs", "confidence threshold", "ECE", asks how to get a
  confidence score out of a model, or asks whether to replace a classification
  prompt with something cheaper.

  DO NOT USE FOR: prose generation, code generation, multi-step reasoning, or
  anything needing a written justification — those are the cases these models
  explicitly cannot serve. For SDK/API code load `typesafe-jev`; for measuring
  and fitting probabilities load `decision-model-calibration`.
allowed-tools: Read, Grep, Glob, Write, Edit, WebFetch
---

# Typed decision models

A class of model that does not generate text. You declare the answer space up
front; the model distributes probability inside it. TypeSafe's **Jev** is the
first sold under that banner ("System One", after Kahneman's fast thinking —
launched in early access on 2026-09-15), and Pydantic AI already abstracts the
class behind a `DecisionModel` base so other backends can follow.

The point is not speed. The point is that **the answer and its uncertainty come
back as separate things**, so code can act on one and route on the other.

> Facts below were verified against primary sources on **2026-10-03**. This
> vendor moves fast: its limitations page was rewritten on 2026-10-02 and its
> rate limit doubled the same week. Re-check anything operational on
> `docs.typesafe.ai` before quoting it.

## Related skills

| Need | Load |
|---|---|
| Python/JS SDK, HTTP API, limits, pricing, framework integrations | `ai-integration/typesafe-jev` |
| Building the labelled corpus, ECE, temperature/Platt fits, option-order debiasing, conformal sets, thresholds | `ai-integration/decision-model-calibration` |

## The three primitives

| | asks | returns | limits (jev-1.13) |
|---|---|---|---|
| **Noul** | is this true? | `noul`: P(yes), 0–1 | no `confidence` field — 0.5 *is* "unsure" |
| **Choice** | which of these? | `choice`, `probabilities` (sum to 1), `confidence` | ≤ 255 options |
| **Score** | at what level? | `score` (probability-weighted, lands *between* levels), `legend`, `probabilities`, `confidence` | 2–10 ordered levels |

Questions in one request share a single read of the `state` and are
**independent**: one answer never becomes context for another. A decision that
depends on a previous one needs a second call. Batching is the main economic
lever — the vendor's own cookbook measured 13 questions in one call at 12.2×
cheaper and 10.0× faster than 13 separate calls.

## What `confidence` is — and is not

`confidence` is a **dispersion statistic computed from `probabilities`**, not a
probability of being right:

- Choice: `(p_max − 1/n) / (1 − 1/n)` — so `(0.6, 0.3, 0.1)` and `(0.6, 0.2, 0.2)` both give 0.4.
- Score: `max(0, 1 − Σ pᵢ·|i − peak| / MAD_uniform)` — mass on a neighbouring level costs less than mass far away.
- Noul: none returned; the vendor suggests `|2p − 1|` if you need one on the same scale.

Pydantic AI says it outright: *"It is a margin, not a probability that the
answer is right."* Consequences:

- A threshold on `confidence` is an **ordering** until you have checked it against
  labels. Only after calibration does "0.9" mean "right 9 times in 10".
- Thresholds **do not transfer** between question kinds (a bar tuned on a Noul
  is meaningless on a Choice), between fields, between backends, or between
  model versions. Pin the version (`jev-1.13.0`, not `jev-latest`) once tuned.

## The shape that actually works

Keep every question atomic — glance work, not reasoning. A complex judgement is
decomposed into several questions and recomposed **in code**, where weights
change under code review instead of inside a prompt. The vendor calls this
*"probably the most important concept in this guide."*

The best public demonstration is a community benchmark on 2,000 phishing emails
(anisselbd/jev-phishing-bench, 2026-09-17): one broad question scored **62.6%**;
five atomic Noul signals combined by a logistic regression fitted on 1,000
labelled emails scored **95.0%** on the other 1,000. Read the caveats the repo
itself states before quoting it: Claude Haiku 4.5 asked the *same five
questions* reached 93.2% (difference not significant, p = 0.063), and a
two-feature regex reached 91.8% because the dataset "largely separates by
construction". **The decomposition did the work — for both models.**

```python
from typesafe_sdk import Choice, Noul, Score, TypeSafeClient

client = TypeSafeClient()  # TYPESAFE_API_KEY from the environment
result = client.system_one(
    {"ticket": ticket},                      # state: string, object or array
    {
        "department":  Choice(instructions="Which team should handle `ticket`?",
                              criteria={"billing": "Payments, invoicing, refunds",
                                        "technical": "Bugs, outages, integrations",
                                        "other": "Anything else"}),
        "is_urgent":   Noul(instructions="Does `ticket` convey urgency?"),
        "frustration": Score(instructions="How frustrated is the customer?",
                             criteria=["Calm, just stating facts",
                                       "Frustrated but civil",
                                       "Very angry, strong language"]),
    },
)
```

Gate on confidence with bars set by the **cost of being wrong**, per action —
the vendor's own example uses a 0.6 floor and 0.85 for a money-moving action:

```python
action = result.choices["department"]
if action.confidence < 0.6:
    route_to_human()
elif action.choice == "technical":
    open_bug()              # cheap to undo: the floor is enough
elif action.confidence > 0.85:
    act_automatically()
else:
    ask_first()
```

Those numbers are starting points the vendor labels as such ("Start with
conservative thresholds, test with your own data") — never production values.

## What the evidence says, as of 2026-10-03

Read this before quoting a vendor number at anyone.

**The headline accuracy is agreement, not correctness.** TypeSafe's 67.8% on
evals.typesafe.ai is agreement with the average of two frontier models at high
thinking, on four workflows it chose; its own blog admits this *"biases answers
towards OpenAI and Anthropic's models."* No human ground truth.

**The "0% hallucination" bar is not measured.** Verbatim footnote: *"Our number
is not empirical. Schema matching is guaranteed, thus we can confidently add 0%
into the plots."* The guarantee is structural. **It cannot produce an invalid
shape; it can absolutely decide wrong.**

**Independent accuracy: mid-tier, not frontier, much cheaper.**
- arXiv 2609.24574 (18 annotation tasks, 7,977 items): Jev trails the best LLM on
  14 of 15 tasks, median −11.6 macro-F1, at a median **44× lower cost**. Routing
  low-confidence items to an LLM **matched the LLM alone at ¼–½ of its cost** —
  the strongest published argument for the composition below.
- BANKING77: eight independent runs ranged 0.753–0.840 (median 0.809). The
  frequently quoted 75.3% is the *lowest* — n = 77, one item per intent.
- Latency: one pre-registered eval measured median calls ~2.2× faster than
  gpt-5.4-nano, not the 40–200× of the launch post (whose own blog says its
  numbers are "on the higher end of real world gains").

**Calibration: decent for its class, not trustworthy raw.**
- Better than the verbalized confidence of 16 of 19 LLMs tested in arXiv
  2609.24574; three frontier models still did better.
- Distortion reported as **compression toward the middle**; probabilities are
  rounded to 0.01 and Choice/Score can return exact 0 or 1.
- A 900-ticket synthetic study measured ECE 0.107 (4.4× its noise floor). Its
  author **corrected** the fitted temperatures on 2026-09-22 (Choice 3.29 → 1.30,
  Score 3.40 → 1.92) after finding that flooring exact zeros at 1e-6 inflated
  them — "read the sign, not the magnitude". Its worst case (44.7% correct at
  mean probability 0.74) was a question **unanswerable from the text by design**:
  the finding is that Jev did not lower its confidence when it could not know.
- Fitting is cheap and effective: a single temperature cut ECE ~74% in one
  review; a Platt intercept on 50 labels cut it 62% in another — while **fewer
  than ~30 labels could make it worse**.

**Option order and option names move the answer.** Since 2026-10-02 the vendor
documents it: *"the order of a Choice's options can affect the answer, and
`jev-1.13` leans toward the option that comes first."* Measured effects:
- a reference card's position moved mean probability on the right answer from
  0.50 (first) to 0.89 (last) (Archer Hume, 10,000 calls);
- adding an irrelevant fifth option shrank the log-odds between two others from
  +0.49 to +0.08 — options **interact**, so probabilities are not fixed
  per-option scores (Hume declines to name the mechanism);
- renaming labels is worse than reordering them: arXiv 2609.26758 reports AUC
  0.81 → 0.58 on the hosted model with 24× the test-retest flip rate.

**Not deterministic.** The vendor's own self-consistency cookbook: picked labels
*"can flip inside a single condition, including TypeSafe… TypeSafe flips on 2
of the 8 questions."* Pydantic AI: numbers "move by a few hundredths from one
run to the next, so a bar is a range to choose from rather than a point."

**Not English? Expect a cost.** Two pre-registered audits exist:
- Russian (n = 600 paired): XNLI accuracy 88.3% → 77.3% and ECE 0.032 → 0.096;
  MASSIVE showed no detectable difference.
- Spanish (19,200 calls): −3.0 to −6.4 pp across XNLI, PAWS-X, MASSIVE and
  Belebele, ECE roughly doubled on XNLI; writing the instructions in Spanish
  did not help.
- No Italian evaluation was found. On any non-English corpus, measure first.

**Limits the vendor documents itself** (jev-1.13 jaggedness page): literal
reading; not a calculator and *"does not count reliably"*; *"reads dates as text,
not as ordered quantities"*; double negatives and indirection; accuracy falls
as the state grows with irrelevant content (*"context rot"*); **state is not
treated as hostile — prompt injection through the state is live**;
contradictory instructions and criteria; option order; not trained to
generate.

## Getting a probability without Jev

**Anthropic cannot give you one.** The Messages API has no `logprobs`
parameter; on the OpenAI-compatibility layer `logprobs` and `top_logprobs` are
*"Ignored"* and the response field *"Always empty."* Claude can be a **typed**
classifier — structured outputs support `enum` — but never a **probabilistic**
one. Confidence from Claude means sampling frequency, verbalized confidence, or
a panel of readers voting, each costing more than one call.

Where logprobs exist (OpenAI `top_logprobs` 0–20 but only with reasoning effort
`none` on GPT-6; Gemini `responseLogprobs` + `logprobs` 0–20; vLLM; llama.cpp
`n_probs`), constrain the label to a **single token** in an enum so that one
token's distribution *is* the class posterior, then calibrate on your labels.
Know which distribution you are reading: vLLM's default `logprobs_mode` is
`raw_logprobs`, *before* logit processors — so a constraint is not reflected in
it. The provider matrix is in `decision-model-calibration`.

## When to reach for this, and when not to

**Yes**: the decision repeats at volume, the answers are known in advance, and
seconds or cents matter. Routing, triage, moderation, tool selection, scoring
large tables, re-ranking retrieved passages, gating a tool call before it runs.

**No**: you need generated text, a written justification, multi-step reasoning,
arithmetic, date comparison, or the problem is one of a kind. And **no** where an
audit needs a reason: a number without a *why* is a defect in regulated
contexts, not a detail.

The honest use is the composition: **the typed model decides fast; the LLM
takes whatever it declares uncertain.** Watch the hand-off rate — a chain that
hands off most requests costs an LLM call *plus* a decision call and is slower
than no decision model at all.

## The cheapest test that can change your mind

Before wiring anything: take a corpus you have already labelled — your own
language, your own state shape — send it through the candidate, and compare
accuracy, calibration and hand-off rate. At $0.042 per million input tokens a
few hundred items cost cents. An afternoon settles what a launch post cannot.

Attribution

claude-dev-suiteclaude-dev-suite
View sourceSee grades on GitHubMore from claude-dev-suite →
SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments (0)

No comments yet. Be the first to comment!

SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Related Skills

Caveman

Terse caveman voice: answer first, fluff gone, every technical fact kept. Use for /caveman, "caveman mode", "talk like caveman", "be brief", "less tokens". Stays on until "stop caveman" or "normal mode".

1100021 votes

Hyperplan

Adversarial multi-agent planning skill. Self-orchestrates 5 hostile category members (unspecified-low, unspecified-high, deep, ultrabrain, artistry) via team-mode for ruthless cross-critique debate, distills only the defensible insights, then MANDATORILY hands the distilled insight bundle to the `plan` agent for executable plan formalization. Use when planning needs maximum rigor and surfacing of weak assumptions, blind spots, and over-engineering. Triggers: 'hyperplan', 'hpp', '/hyperplan', ...

698431 votes

Writing Skills

Create and manage Claude Code skills in HASH repository following Anthropic best practices. Use when creating new skills, modifying skill-rules.json, understanding trigger patterns, working with hooks, debugging skill activation, or implementing progressive disclosure. Covers skill structure, YAML frontmatter, trigger types (keywords, intent patterns), UserPromptSubmit hook, and the 500-line rule. Includes validation and debugging with SKILL_DEBUG. Examples include rust-error-stack, cargo-dep...

3931 votes

Mcp Code Execution

Routes multi-tool workflows through MCP servers for large datasets and pipelines. Use when Bash tool overhead is limiting throughput on data-heavy tasks.

3421 votes

catchup

Recovers the conversation and failed tool calls of a previous Codex, Amp, Claude Code, Antigravity, Cline, Copilot CLI, Cursor, DeepSeek Harness, Grok Build, Kimi, OpenCode, Pi Agent, or ZCode session. Use when the user says "catch up", "what did the last session do", "get me up to speed", "I switched agents", asks to recover/summarize a previous session before continuing, or asks to diagnose or report a catchup failure. Do NOT use for the current conversation, git history, or any non-agent log.

741 votes
View all in ai-agents →