Learn whether a fine-tuned model is actually better than what it started from, instead of guessing from a few prompts that felt good. Builds a probe set from the task definition, writes a rubric, runs the tuned model against its own base model and optionally against a hosted frontier model, scores the results, and produces a scorecard you can rerun after every training run. Use when someone asks whether a fine-tune worked, whether it is good enough to ship, how it compares to the base model o...
Installs into .claude/skills of the current project.
Are you the author of Evaluating A Tuned Model?
Add the live security badge to your README. It updates with every re-scan.
[](https://www.skillsdirectory.com/skills/ertasai-evaluating-a-tuned-model)
---
name: evaluating-a-tuned-model
description: >-
Learn whether a fine-tuned model is actually better than what it started
from, instead of guessing from a few prompts that felt good. Builds a probe
set from the task definition, writes a rubric, runs the tuned model against
its own base model and optionally against a hosted frontier model, scores the
results, and produces a scorecard you can rerun after every training run. Use
when someone asks whether a fine-tune worked, whether it is good enough to
ship, how it compares to the base model or to an API, or wants a regression
suite for future training runs. Not for fixing a model that behaves badly at
inference, which is debugging-a-bad-fine-tune, and not for measuring cost,
which is costing-a-model-vs-an-api.
license: Apache-2.0
metadata:
version: "0.1.0"
author: "Edward Xi Yang, Ertas AI"
stage: "evaluate"
previous_skill: "inspecting-a-model-bundle"
next_skill: "costing-a-model-vs-an-api"
---
# Evaluating a tuned model
## Compare it against its own base, not against a vague memory of "before"
This is the argument the rest of this skill exists to make operational, so
it comes first: a fine-tune has to be measured against the base model it
started from, run on the same probes, under the same decoding settings.
Most claimed improvements do not survive that comparison. A handful of
prompts that felt better after training, a training loss curve that went
down, or a memory of how bad the base model used to be are not a
measurement, they are an impression. Skip the base comparison and ship on
an impression, and the first time someone asks "how do you know it's
better" there is no answer.
The base model is also the only fair opponent. Comparing a fine-tune to a
much larger frontier model and losing tells you nothing about whether
training worked, because you already knew a small model would lose to a
frontier model before training even started. The question training answers
is narrower and more useful: did this task-specific training move the
needle from where this exact base model started. A frontier comparison is
still worth running, but as a second, optional data point about where the
tuned model sits in the wider market, not as the test of whether training
did anything.
## The shape of a defensible eval
Four pieces, produced in this order:
1. **`eval/probes.jsonl`**, a set of held-out inputs the model has to
handle, one JSON object per line, each shaped
`{"id": "...", "prompt": "...", "expected_behaviour": "..."}`.
2. **`eval/rubric.md`**, the pass and fail conditions for those probes,
written so a human or a judge model applies it the same way twice.
3. **A results file**, one line per probe, recording which model won that
probe: `{"id": "...", "winner": "tuned" | "base" | "tie"}`. This is
produced by running the rubric against paired outputs from the base and
tuned models, not by this skill's script.
4. **`eval/scorecard.json`**, produced by scoring the results file with
`scripts/score_eval.py`.
If `BUNDLE-REPORT.md` exists in the project root from
**inspecting-a-model-bundle**, read its **Base model** section first. It
names exactly which base checkpoint the tuned model was trained from, which
is what "base" has to mean in every arm of this comparison. If that file
does not exist, get the base model identifier directly: for an adapter,
`adapter_config.json`'s `base_model_name_or_path`; for a merged checkpoint,
whatever the training run recorded; ask rather than guess if neither is
available, because comparing against the wrong base produces a number that
looks precise and means nothing.
## Build the probe set from the task, not from the training data
Write probes that describe what the model has to be able to do, then check
whether each training example happens to already be one. Do not go the
other direction and pull probes out of the training set after the fact.
Probes sourced from training data tend to measure memorisation of that
exact phrasing rather than the underlying task, and inflate the tuned
model's apparent win rate for a reason that has nothing to do with whether
it generalises.
Hold out real examples the model has never seen: inputs from the same
distribution as production traffic or real user requests, set aside before
training starts, not synthesised afterward to fit whatever the model
happens to already do well. Cover the behaviours that actually matter for
the task, not just the easy middle of the distribution, and include a
handful of adversarial or edge cases deliberately, since those are usually
where a fine-tune quietly regresses even while the easy cases look fine.
Full detail on deriving probes from a task definition, coverage across
behaviours, held-out examples, and why fewer than 20 probes is noise rather
than a measurement, is in `references/building-probes.md`.
## Write a rubric before looking at any output
Write the rubric against the task definition, before either model has
generated anything for you to react to. A rubric written after seeing the
outputs tends to describe what the tuned model happened to do well, which
makes the comparison circular. For each probe, or each class of probe, the
rubric should state a concrete pass condition, not "good" or "sounds
right": does it contain the required fields, does it follow the requested
format, does it avoid a specific known failure mode, does it match a
reference answer for probes that have one.
`references/judging-and-scoring.md` covers writing the rubric itself,
scoring by a human versus a judge model, the specific biases judge models
bring to this task, and how to blind a comparison so those biases matter
less.
**Where a judge model applies the rubric, control it before you trust it.**
Run a handful of probes whose answers you already know are good against
answers you know are bad, and check the judge prefers the good ones. Run the
same output against itself and check it comes back a tie. A judge that fails
either check still produces verdicts, and those verdicts will look exactly
like real ones on the scorecard. The procedure is in
`references/judging-and-scoring.md`.
## Run the comparison: base, tuned, and optionally a frontier model
Run every probe against the base model and the tuned model with identical
prompts and identical decoding settings (temperature, max tokens, stop
sequences). If either arm samples rather than decodes greedily, run it more
than once per probe and treat the outcome as noisy rather than a single
verdict.
Judge each pair against the rubric and record one line per probe in a
results file:
```json
{"id": "p014", "winner": "tuned"}
```
`winner` is `"tuned"`, `"base"`, or `"tie"`. A tie is a real, useful
outcome, not a failure to decide: it means the rubric could not separate
the two outputs on that probe, and that is worth recording exactly as much
as a win.
### Optional: compare against a hosted frontier model
A third arm, a hosted frontier model answering the same probes, tells you
where the tuned model sits against the wider market rather than just
against where it started. This section is optional and requires your own
API key for whichever provider you use. No script in this skill calls out
to a network, and nothing above this point needs one; running this arm is
your decision to make with your own credentials, not something this skill
does on your behalf.
If you run it, judge the frontier output against the same rubric and add it
to the results as a separate probe-provider pair, or run a second results
file for the frontier-vs-tuned comparison and score it separately. Keep the
base-vs-tuned comparison as the one that answers "did training work"; treat
the frontier comparison as an additional, clearly labelled data point, not
a replacement for it.
## Score it
From this skill's own directory:
```bash
python3 scripts/score_eval.py eval/results.jsonl > eval/scorecard.json
```
Python 3.9 or newer, standard library only, no network access.
The scorecard's `totals` counts `tuned`, `base`, and `tie`, plus `decided`
(`tuned` + `base`) and `n` (every row). `win_rate_vs_base` is the tuned
model's share of `decided`, not of `n`. **Ties are excluded from that
denominator on purpose.** Counting a tie as a decided comparison in either
direction manufactures a result out of "the rubric could not tell them
apart", which is a different fact from a win or a loss and would quietly
inflate or deflate the headline number depending on which way you folded
it in. If more than half the probes come back tied, the scorecard's
`warnings` will say so directly: that usually means the rubric is not
strict enough to discriminate between the two models, not that the models
are equivalent.
The scorecard also carries `warnings` for too few probes, and a plain-text
`verdict` that states the result honestly, including when the result is a
wash. A verdict is never phrased as an improvement unless the win rate over
decided comparisons clears a real margin. At close to 50%, the honest
reading is that training did not move the needle on this probe set, and the
scorecard says exactly that rather than rounding a wash up to a win.
**If Python is unavailable, or the script errors,** score by hand: open the
results file, count how many rows have `winner: "tuned"`, `"base"`, and
`"tie"`, and compute `tuned / (tuned + base)` for the win rate. This is the
entire computation the script does; nothing about it needs the script to
be correct, only to be faster than counting by hand on a large probe set.
## Read the scorecard honestly
A tuned model that beats its base by a wide, sustained margin across a
probe set that actually covers the task is a real result. A tuned model
that ties or narrowly edges out its base is not evidence training was
wasted, but it is evidence the current training run did not move this
particular probe set, and that is worth saying plainly rather than
stretching a 52% win rate into a headline. A tuned model that loses to its
base is the most important result of the three to report accurately: it
means the training run made the model worse at this task, and shipping it
anyway on the strength of a good-looking training loss curve is exactly
the failure this skill exists to catch.
If the result is a loss or a wash and the generated outputs themselves look
wrong, not just unimproved, that is a different problem: see
**debugging-a-bad-fine-tune**.
## Rerun this as a regression suite
Version `eval/probes.jsonl` and `eval/rubric.md` alongside the model
config, and rerun the same probe set after every training run, not just
the first one. A model that improved on the first fine-tune and regressed
on the second is a common outcome, and the only way to see it is to keep
comparing against the same fixed set rather than re-deriving a fresh, more
flattering probe set each time. When the task itself changes enough that
the probe set no longer reflects it, version the probe set explicitly (a
date or a number in the filename is enough) so an old scorecard is never
silently compared against a new probe set as if they measured the same
thing.
## Hand off to
- The generated output itself looks broken, not just unimproved: **debugging-a-bad-fine-tune**
- The tuned model clears the quality bar and the next question is whether
it is worth running instead of an API at your usage level: **costing-a-model-vs-an-api**
- The bundle's shape or base model was never confirmed before running this
comparison: **inspecting-a-model-bundle**