Skip to content
Back to skills

Run

ASecurity

Run the agent-config eval suite locally and compare with the stored baseline; explain why a case went red. Use when the user says "run the evals", "run config-drift-checker", "did the Claude Code update break my setup", "why is the eval red", or "compare against baseline". Do not use to change the setup itself (that is repair) or to write new cases (that is write-case).

  • 30 stars
  • 0 votes
  • 0 copies
  • 0 views
  • Added October 10, 2026
ai-agentsrustbashnodegit

Works with

  • claude code

Security analysis

A100/100

Scanned October 10, 2026

npx -y skills add jameskomo/config-drift-checker --skill run --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Run?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Run
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/jameskomo-run/badge)](https://www.skillsdirectory.com/skills/jameskomo-run)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: run
description: Run the agent-config eval suite locally and compare with the stored baseline; explain why a case went red. Use when the user says "run the evals", "run config-drift-checker", "did the Claude Code update break my setup", "why is the eval red", or "compare against baseline". Do not use to change the setup itself (that is repair) or to write new cases (that is write-case).
---

# config-drift-checker: run

```
node ${CLAUDE_PLUGIN_ROOT}/tools/eval-shim.mjs <plugin> [--track pinned|canary] [--case <glob>] [--runs n] [--ablation none|with-without] [--scaffold] [--model m] [--budget <usd>]
node ${CLAUDE_PLUGIN_ROOT}/tools/eval-diff.mjs <baseline.json> <current.json> --config <plugin>
node ${CLAUDE_PLUGIN_ROOT}/tools/config-coverage.mjs <plugin>              # which rules have no case
node ${CLAUDE_PLUGIN_ROOT}/tools/release-watch.mjs --state .release-watch.json --models --pin <model.pinned>
```

- Default to `--ablation none --scaffold` for regression checks (with-without only to prove a
  plugin's worth). Use `--runs 1` for a quick look, 3+ before trusting a score. Always pass a
  `--budget` when the user has said what they are willing to spend; the `.cdc.yml` `budget.per_run_usd`
  applies otherwise.
- `--track pinned` = the baseline (exact model id + Claude Code version from `.cdc.yml`);
  `--track canary` = the alias on whatever Claude Code is installed, 1 run per case, more only on a
  deviation. "Did the update break my setup?" is a canary question; "did my edit break it?" is pinned.
- Baseline lives at `<plugin>/evals/results/<timestamp>/aggregate-result.json` locally, or
  `baseline.json` on the `eval-results` branch (`git show origin/eval-results:baseline.json`).
- The diff header says whether the **model** or **Claude Code** moved since the baseline, and flags
  *slower / pricier / longer* cases (efficiency drift) even when every case still passes.
- If the official runner is enabled (`claude plugin eval` in an empty dir prints "No eval cases
  found"), prefer `claude plugin eval <plugin> --allow-tools Bash --scaffold --json out.json`.

## Explaining red: read the run, not the score

For each failing case open the run entries in the JSON: `numTurns`, `toolUses`, `response`,
per-grader `verdict`. Classify and say which:
1. **Model refused before acting** (1 turn, 0 tool calls) → the case doesn't exercise the hook; change the command/scaffold.
2. **Hook/skill didn't fire** (tool attempted, no block / no Skill call) → real regression or config change; the report's stamp says whether the model or Claude Code moved. Offer `/config-drift-checker:repair`.
3. **Grader wrong** (prose matched, negative grader with min=1) → fix the grader, not the setup.
4. **Flaky** (mixed verdicts across runs) → raise `runs`; never loosen the threshold.

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…