Run the golden-case eval suite against the current branch,
Scanned 9/19/2026
Install to Claude Code
npx -y skills add thecoderpanda/fde-starter-kit --skill agent-eval --agent claude-codeInstalls into .claude/skills of the current project.
Are you the author of Agent Eval?
Add the live security badge to your README — it updates automatically with every re-scan.
[](https://www.skillsdirectory.com/skills/thecoderpanda-agent-eval)More formats (shields.io, HTML) on the badges page.
---
name: agent-eval
description: Run the golden-case eval suite against the current branch,
compare pass rates to main, and summarize what regressed and why. Use
after any change to files under ./agent/, ./app/api/chat/, or ./evals/.
---
# agent-eval
Purpose: turn "does this change break the agent?" from a vibe check into
a diffable report in under 90 seconds.
## When to fire
- The user says any of: "run evals", "check evals", "did I break
anything", "eval my change", "compare against main".
- The user just edited a prompt, a tool schema, a tool's execute body,
the system prompt, `./agent/config.ts`, or added / changed a case in
`./evals/cases/`.
## Preconditions
Before running the suite, verify:
1. `OPENAI_API_KEY` is set (`printenv OPENAI_API_KEY | head -c 8`).
2. `node_modules/` exists — run `npm install` if not.
3. `npm run typecheck` passes. A type error in a tool schema will make
every case fail with a confusing message; catch it upfront.
If any precondition fails, stop and report — do not run the suite.
## Steps
1. **Snapshot the current branch's results.**
```bash
npm run eval:ci
```
The runner writes `eval-results/summary.json`. Keep this file's path
handy for step 4.
2. **Snapshot main's results for comparison.**
```bash
git stash push -u -m "agent-eval:pre-main"
git checkout main -- .
npm run eval:ci
cp eval-results/summary.json eval-results/summary.main.json
git checkout HEAD -- . # restore working tree
git stash pop
```
If the user is already on `main`, skip this step and note that no
comparison baseline is available.
3. **Diff the two runs.** For each case ID present in both files,
compare `passed`. Bucket into:
- **New failures** — passed on main, failed on branch. These are
regressions you should call out first.
- **New passes** — failed on main, passed on branch. Improvements.
- **Consistent failures** — failed on both. Pre-existing debt.
- **Consistent passes** — the boring majority.
4. **Report.** Write a short summary in this shape:
```
Eval delta vs main
Branch: <passed>/<total>
Main: <passed>/<total>
New failures (N):
- <case.id> <one-line why, quoting the first failure message>
New passes (N):
- <case.id>
Consistent failures still open (N):
- <case.id>
```
Do NOT paste the full JSON. Do NOT include cases that didn't change.
5. **Hypothesize.** For each new failure, read the case in
`./evals/cases/*.jsonl`, then read the file(s) the user just changed,
and offer one specific hypothesis for the regression. Do not guess if
the case is unfamiliar — say so.
## Definition of done
- Both runs completed (or you explicitly reported the missing baseline).
- The report exists and lists new failures first.
- Each new failure has a hypothesis or an explicit "unknown, needs
investigation" note.
- You did NOT commit anything, edit prompts, or "fix" failures without
asking the user first.
Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.
No comments yet. Be the first to comment!