Skip to content
Back to skills

Decision Model Calibration

ASecurity

Turning a model's class probabilities into numbers you can set a threshold on: building the labelled corpus, measuring accuracy and calibration honestly (ECE with a noise floor, Brier, risk–coverage), fitting temperature/Platt, removing option-order bias, conformal prediction sets, and choosing per-action thresholds. Applies to Jev and any typed classifier with probabilities, and says which LLM providers expose logprobs at all. USE WHEN: user asks whether a confidence/probability can be trus...

  • 40 stars
  • 0 votes
  • 0 copies
  • 0 views
  • Added October 3, 2026
ai-agentspythonrustgobashgitapi

Works with

  • cli
  • api

Security analysis

A100/100

Scanned October 3, 2026

npx -y skills add claude-dev-suite/claude-dev-suite --skill decision-model-calibration --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Decision Model Calibration?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Decision Model Calibration
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/claude-dev-suite-decision-model-calibration/badge)](https://www.skillsdirectory.com/skills/claude-dev-suite-decision-model-calibration)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: decision-model-calibration
description: |
  Turning a model's class probabilities into numbers you can set a threshold on:
  building the labelled corpus, measuring accuracy and calibration honestly (ECE
  with a noise floor, Brier, risk–coverage), fitting temperature/Platt, removing
  option-order bias, conformal prediction sets, and choosing per-action
  thresholds. Applies to Jev and any typed classifier with probabilities, and
  says which LLM providers expose logprobs at all.

  USE WHEN: user asks whether a confidence/probability can be trusted, how to
  pick a threshold, about "calibration", "ECE", "reliability diagram",
  "temperature scaling", "Platt", "isotonic", "CalibratedClassifierCV",
  "conformal prediction", "abstain", "selective prediction", "risk coverage",
  "position bias", "option order", "logprobs", or wants to evaluate Jev or an
  LLM classifier against labels.

  DO NOT USE FOR: training classifiers from scratch, or generic LLM evals of
  free-text output (no class probabilities to calibrate).
allowed-tools: Read, Grep, Glob, Write, Edit, Bash
---

# Calibrating decision models

A probability is useful only if "0.9" means right about nine times in ten **on
your data**. No vendor can promise that for you: calibration depends on your
inputs, your language and your label set, and it shifts with every model
version. This skill is the procedure; `typed-decision-models` has the evidence
about Jev in particular. The code below was run on simulated data with known
over-confidence before it was written here (Python 3.12, numpy 2.5, scipy 1.18,
scikit-learn 1.9.1).

## 1. The corpus comes first

It outlives every model you try against it, and it is the only thing that lets
you compare them.

- **Draw from production traffic**, not from what is easy to label. Same
  language, same state shape, same class balance — and include the ambiguous
  cases, which are the ones thresholds exist for.
- **Size.** About 50 labels per question can fit a one-parameter correction;
  **fewer than ~30 can make calibration worse** (observed on Jev). Isotonic
  regression needs ~1,000 (scikit-learn: "not recommended when calibration
  samples < ~1000"). Thresholds for rare, expensive classes need enough
  examples *of that class*.
- **Split** into a calibration half and a test half; never report a metric on
  the data a correction was fitted on.
- **Label what a person on the team would do**, not what the model said. A
  corpus labelled by another model measures agreement, not correctness — the
  exact flaw in the Jev launch benchmark.
- **Record the model version** alongside every stored probability, so a
  re-tune can be triggered when it changes.

## 2. Measure

```python
import numpy as np


def ece(confidence, correct, n_bins=10):
    """Expected calibration error with equal-mass bins (less biased than equal-width)."""
    confidence = np.asarray(confidence, float)
    correct = np.asarray(correct, float)
    order = np.argsort(confidence)
    total = len(confidence)
    return sum(
        len(b) / total * abs(correct[b].mean() - confidence[b].mean())
        for b in np.array_split(order, n_bins)
        if len(b)
    )


def ece_noise_floor(confidence, n_bins=10, n_sims=2000, seed=0):
    """ECE a *perfectly calibrated* model would show on this many items with these
    confidences. A measured ECE inside this band is indistinguishable from noise."""
    rng = np.random.default_rng(seed)
    confidence = np.asarray(confidence, float)
    sims = [ece(confidence, rng.random(confidence.size) < confidence, n_bins) for _ in range(n_sims)]
    return np.percentile(sims, [50, 95])


def brier(probs, labels):
    """Multiclass Brier score; probs is (n, k), labels are column indices."""
    probs = np.asarray(probs, float)
    onehot = np.eye(probs.shape[1])[labels]
    return np.mean(np.sum((probs - onehot) ** 2, axis=1))
```

Report, per question:
- **accuracy with a confidence interval**. n = 77 gives roughly ±10 points;
- **ECE next to its noise floor**. At n = 60 a perfectly calibrated model already
  scores ≈ 0.045, so a bare ECE is uninterpretable;
- **Brier score**, which punishes both miscalibration and poor discrimination;
- **direction**: is the error over-confident (most binned accuracy below binned
  confidence) or compressed toward the middle?

What "confidence" means:
- **Choice:** use `p_max`, or the probability of the chosen option.
- **Noul:** use `max(p, 1 − p)`, with correctness judged at your threshold.

**Do not calibrate Jev's `confidence` field.** It is a dispersion statistic
derived from the distribution (`(p_max − 1/n)/(1 − 1/n)` for Choice), not a
probability. Calibrate `probabilities` or `noul`.

## 3. Fit

### Choice / Score: one temperature

```python
from scipy.optimize import minimize_scalar
from scipy.special import softmax

# Jev rounds probabilities to 0.01 and returns exact zeros. log(0) is -inf, and a
# tiny floor (1e-6) turns every zero into a huge negative logit that inflates the
# fitted temperature. Half the rounding quantum is the honest floor.
FLOOR = 0.005


def to_matrix(answers, options):
    """[{'billing': 0.88, ...}, ...] -> (n, k) array in a fixed option order."""
    return np.array([[a[o] for o in options] for a in answers], float)


def fit_temperature(probs, labels, floor=FLOOR):
    logits = np.log(np.clip(probs, floor, 1.0))
    labels = np.asarray(labels)

    def nll(t):
        p = softmax(logits / t, axis=1)
        return -np.mean(np.log(p[np.arange(len(labels)), labels]))

    return minimize_scalar(nll, bounds=(0.05, 20.0), method="bounded").x


def apply_temperature(probs, t, floor=FLOOR):
    return softmax(np.log(np.clip(probs, floor, 1.0)) / t, axis=1)
```

T > 1 means the model was over-confident; T < 1 means it was under-confident.
The floor is not a detail:
- **Simulation:** true over-confidence T = 2.0, outputs rounded to 0.01 like
  Jev's. A 0.005 floor recovered T = 1.84; a 1e-6 floor reported **3.02**.
- **Published study:** the same artefact made an independent Jev audit publish
  temperatures of 3.29 and 3.40. Its author corrected them on 2026-09-22 to
  1.30 and 1.92.

### Noul: Platt (two parameters)

```python
from scipy.special import logit
from sklearn.linear_model import LogisticRegression


def fit_platt(p_yes, y, floor=FLOOR):
    x = logit(np.clip(np.asarray(p_yes, float), floor, 1 - floor)).reshape(-1, 1)
    return LogisticRegression(C=1e6).fit(x, y)  # effectively unregularised: two parameters


def apply_platt(model, p_yes, floor=FLOOR):
    x = logit(np.clip(np.asarray(p_yes, float), floor, 1 - floor)).reshape(-1, 1)
    return model.predict_proba(x)[:, 1]
```

Platt's slope also corrects **compression toward 0.5**, which independent work
reports for Jev and which a pure shift cannot fix.

### Your own scikit-learn classifier

scikit-learn ≥ 1.8 has `method="temperature"` alongside `"sigmoid"` and
`"isotonic"`. `cv="prefit"` was deprecated in 1.6; wrap an already-fitted model
in `FrozenEstimator`:

```python
from sklearn.calibration import CalibratedClassifierCV
from sklearn.frozen import FrozenEstimator

calibrated = CalibratedClassifierCV(FrozenEstimator(clf), method="temperature").fit(X_cal, y_cal)
```

### What one fit buys

- **Classic result:** temperature scaling alone fixes most of the
  miscalibration of modern networks (Guo et al., ICML 2017).
- **On Jev:** a single temperature cut ECE by ~74% in one review; a Platt
  intercept fitted on 50 labels cut it by 62% in another.
- **This skill's simulation:** ECE went 0.145 → 0.035 on held-out data.

Fit **one correction per question**. Never share one across Noul, Choice and
Score, or across fields.

## 4. Remove option-order bias before trusting a Choice

Multiple-choice selectors favour positions:
- Zheng et al., *Large Language Models Are Not Robust Multiple Choice
  Selectors* (ICLR 2024).
- Pezeshkpour & Hruschka (2023) measured gaps of 13–75% from reordering alone.
- TypeSafe documents that `jev-1.13` "leans toward the option that comes first".

**Cyclic permutation** averages over n orders instead of n!. With Jev, all the
orders go into **one request**: questions are independent and share one read of
the state, so the extra orders cost tokens, not a round trip.

```python
def rotations(criteria, limit=None):
    """Cyclic rotations of a Choice's options (Zheng et al., ICLR 2024): n orders, not n!."""
    items = list(criteria.items())
    n = len(items)
    count = n if limit is None else min(limit, n)
    shifts = sorted({round(j * n / count) % n for j in range(count)})  # evenly spread, first order kept
    return [dict(items[i:] + items[:i]) for i in shifts]


def order_checked_choice(client, state, instructions, criteria, limit=None):
    """Ask the same Choice under several option orders in ONE request and average.

    Questions in a request are independent and share one read of the state, so
    the extra orders cost only their question tokens, not another round trip.
    """
    from typesafe_sdk import Choice

    orders = rotations(criteria, limit)
    questions = {f"order_{i}": Choice(instructions=instructions, criteria=c) for i, c in enumerate(orders)}
    result = client.system_one(state, questions)
    answers = [result.choices[key] for key in questions]
    averaged = {o: sum(a.probabilities[o] for a in answers) / len(answers) for o in criteria}
    winner = max(averaged, key=averaged.get)
    stable = all(a.choice == winner for a in answers)
    return winner, averaged, stable
```

- **Cost:** n orders cost ~n× the question tokens. All of them must fit in the
  64k request budget.
- **Many options:** set `limit` (e.g. 3–4 evenly spaced rotations).
- **How to use it:**
  - Run it **offline over the corpus** first, to measure how often the answer
    flips.
  - Add it in production only for decisions whose `stable=False` rate justifies
    the tokens.
  - Calibrate the *averaged* probabilities, not the single-order ones.
- **Option names matter too.** Renaming labels moved AUC from 0.81 to 0.58 on
  hosted Jev (arXiv 2609.26758). Once a wording is tuned, freeze option names
  like an API.
- **Contextual calibration** (Zhao et al., *Calibrate Before Use*, ICML 2021)
  is the complementary fix for a label prior. Score a content-free input such
  as "N/A", then divide out the bias it reveals.

## 5. Choose thresholds from cost, then check coverage

```python
def risk_coverage(confidence, correct):
    """For each candidate bar: the share of items kept and the error rate among them."""
    confidence = np.asarray(confidence, float)
    correct = np.asarray(correct, bool)
    rows = []
    for bar in np.unique(confidence):
        kept = confidence >= bar
        rows.append((bar, kept.mean(), 1 - correct[kept].mean()))
    return rows


def lowest_bar_for(target_error, confidence, correct, min_kept=30):
    """Lowest bar whose kept items err at most target_error, with at least min_kept items."""
    for bar, coverage, error in risk_coverage(confidence, correct):
        if error <= target_error and coverage * len(confidence) >= min_kept:
            return bar, coverage, error
    return None
```

- **Ask for each direction's cost separately.** For a Noul, the false-positive
  and false-negative costs set where the bar sits. TypeSafe's own advice: raise
  it when acting on a false yes is expensive, lower it when missing a true yes
  is.
- **Read coverage alongside error.** A bar that hands 80% of traffic to an LLM
  costs more than no decision model at all; the hand-off rate is the number
  that shows it.
- **Bars are ranges, not points.** Jev's probabilities move by a few hundredths
  between identical calls, so leave margin.
- **For a guarantee rather than an estimate**, use selective classification
  with a risk bound (Geifman & El-Yaniv, NeurIPS 2017). It picks the bar so
  that the target error holds with high probability.
- **Re-tune whenever any of these change:** model version, option wording,
  option set, or traffic mix.

## 6. Conformal sets: "one of these, with 90% coverage"

When acting needs a guarantee and a single label is too strong, return a
**set**. Use split conformal with the LAC score (Kumar et al. 2023), on
calibrated probabilities:

```python
def conformal_qhat(p_true_label, alpha):
    """p_true_label: calibrated probability each calibration item gave its *true* label."""
    scores = np.sort(1.0 - np.asarray(p_true_label, float))
    n = len(scores)
    k = int(np.ceil((n + 1) * (1 - alpha)))
    return 1.0 if k > n else scores[k - 1]  # too few items: the set must be everything


def prediction_set(probabilities, qhat):
    return [option for option, p in probabilities.items() if 1.0 - p <= qhat]
```

- **What it guarantees:** the true label is in the set at least 1 − α of the
  time, provided traffic is exchangeable with the calibration data. In
  simulation, α = 0.1 gave 91.7% coverage.
- **How to act on it:**
  - singleton → act automatically;
  - two or more options → ask the user or a human;
  - empty → treat as out of distribution.
- **Variants:** APS (Romano et al. 2020) and RAPS (Angelopoulos et al., ICLR
  2021) give adaptive or smaller sets.
- **When the guarantee breaks:** it does not survive a shift in traffic.
  Re-calibrate when the input mix changes.

## 7. Which providers give you probabilities at all (2026-10-03)

| Provider | Class probabilities | Notes |
|---|---|---|
| **TypeSafe Jev** | yes, native | `probabilities` / `noul`, rounded to 0.01 |
| **Anthropic (Claude)** | **no** | Messages API has no logprobs; OpenAI-compat `logprobs`/`top_logprobs` are "Ignored", response field "Always empty". Structured outputs give a typed enum, not a distribution |
| **OpenAI** | yes, restricted | `logprobs` + `top_logprobs` 0–20. On GPT-6, with reasoning effort other than `none` these must be removed — no logprobs from a reasoning model that is reasoning |
| **Google Gemini** | yes | `responseLogprobs: true`, `logprobs` 0–20; `text/x.enum` response type for enums. Check per-model support |
| **vLLM** | yes | `logprobs` (cap `max_logprobs`, default 20); structured `choice` output. Default `logprobs_mode` is `raw_logprobs` — **before** logit processors, so a constraint is not reflected |
| **llama.cpp server** | yes | `n_probs`; `post_sampling_probs` for after the sampling chain; `grammar` / `json_schema` constraints |

With logprobs, constrain the label to a **single token**. For a single-token
enum, a grammar's renormalisation is exactly the conditional class posterior.
For multi-token labels it is not.

**Without logprobs (Claude), each option costs more than one call:**
- **Sample k times** and use label frequency.
- **Verbalized confidence.** It is better calibrated than token probabilities
  for RLHF models (Tian et al., EMNLP 2023), but over-confident (Xiong et al.,
  ICLR 2024).
- **A panel of prompts that vote.**

Post-training makes this worse: GPT-4's MMLU ECE went from 0.007 pre-trained
to 0.074 after RLHF (GPT-4 Technical Report, Fig. 8). Calibrate whatever you
end up with.

## Checklist before a threshold goes to production

- [ ] Labelled corpus from production traffic, in the production language, split cal/test
- [ ] Accuracy with CI, ECE with noise floor, Brier — per question, on the test half
- [ ] Option-order flip rate measured; averaging adopted where it matters
- [ ] One temperature (Choice/Score) or Platt (Noul) fit per question, floor = half the rounding quantum
- [ ] Threshold chosen from asymmetric costs; coverage / hand-off rate recorded
- [ ] Model version pinned and logged; re-tune trigger defined
- [ ] Destructive actions still gated by approval, not by a probability alone

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…