Skills DirectorySkills Directory
SkillsLearnSecurityCategoriesDocsCommunityBlog
Sign InSubmit Skill
Skills Directory

Security-tested agent skills for Claude, coding agents, and AI workflows.

Directory

  • Browse Skills
  • All Skills A–Z
  • Claude Skills
  • Claude Code Skills
  • Agent Skills
  • Categories
  • Submit a Skill

Learn

  • Learn Hub
  • Install Claude Skills
  • Write SKILL.md
  • Skills vs MCP
  • Directories Compared

Security

  • Security
  • Methodology
  • Secure Claude Skills
  • Security Badges

Company

  • About
  • Community
  • Blog
  • API Docs
  • Advertise

2026 Skills Directory. All rights reserved.

Back to skills

Agent Eval

ASecurity

Run the golden-case eval suite against the current branch,

2 stars
0 votes
0 copies
0 views
Added 9/19/2026
ai-agentsgobashnodegitapi

Works with

api

Security Analysis

A92/100
mediumInstalls packages at runtime which could introduce malicious dependencies

Scanned 9/19/2026

Install to Claude Code

$npx -y skills add thecoderpanda/fde-starter-kit --skill agent-eval --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Agent Eval?

Add the live security badge to your README — it updates automatically with every re-scan.

Security grade badge for Agent Eval
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/thecoderpanda-agent-eval/badge)](https://www.skillsdirectory.com/skills/thecoderpanda-agent-eval)

More formats (shields.io, HTML) on the badges page.

Download Zip
Files
SKILL.md
---
name: agent-eval
description: Run the golden-case eval suite against the current branch,
  compare pass rates to main, and summarize what regressed and why. Use
  after any change to files under ./agent/, ./app/api/chat/, or ./evals/.
---

# agent-eval

Purpose: turn "does this change break the agent?" from a vibe check into
a diffable report in under 90 seconds.

## When to fire

- The user says any of: "run evals", "check evals", "did I break
  anything", "eval my change", "compare against main".
- The user just edited a prompt, a tool schema, a tool's execute body,
  the system prompt, `./agent/config.ts`, or added / changed a case in
  `./evals/cases/`.

## Preconditions

Before running the suite, verify:

1. `OPENAI_API_KEY` is set (`printenv OPENAI_API_KEY | head -c 8`).
2. `node_modules/` exists — run `npm install` if not.
3. `npm run typecheck` passes. A type error in a tool schema will make
   every case fail with a confusing message; catch it upfront.

If any precondition fails, stop and report — do not run the suite.

## Steps

1. **Snapshot the current branch's results.**
   ```bash
   npm run eval:ci
   ```
   The runner writes `eval-results/summary.json`. Keep this file's path
   handy for step 4.

2. **Snapshot main's results for comparison.**
   ```bash
   git stash push -u -m "agent-eval:pre-main"
   git checkout main -- .
   npm run eval:ci
   cp eval-results/summary.json eval-results/summary.main.json
   git checkout HEAD -- .   # restore working tree
   git stash pop
   ```
   If the user is already on `main`, skip this step and note that no
   comparison baseline is available.

3. **Diff the two runs.** For each case ID present in both files,
   compare `passed`. Bucket into:
   - **New failures** — passed on main, failed on branch. These are
     regressions you should call out first.
   - **New passes** — failed on main, passed on branch. Improvements.
   - **Consistent failures** — failed on both. Pre-existing debt.
   - **Consistent passes** — the boring majority.

4. **Report.** Write a short summary in this shape:
   ```
   Eval delta vs main
     Branch:  <passed>/<total>
     Main:    <passed>/<total>

   New failures (N):
     - <case.id>  <one-line why, quoting the first failure message>

   New passes (N):
     - <case.id>

   Consistent failures still open (N):
     - <case.id>
   ```
   Do NOT paste the full JSON. Do NOT include cases that didn't change.

5. **Hypothesize.** For each new failure, read the case in
   `./evals/cases/*.jsonl`, then read the file(s) the user just changed,
   and offer one specific hypothesis for the regression. Do not guess if
   the case is unfamiliar — say so.

## Definition of done

- Both runs completed (or you explicitly reported the missing baseline).
- The report exists and lists new failures first.
- Each new failure has a hypothesis or an explicit "unknown, needs
  investigation" note.
- You did NOT commit anything, edit prompts, or "fix" failures without
  asking the user first.

Attribution

thecoderpandathecoderpanda
View sourceMore from thecoderpanda →
SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments (0)

No comments yet. Be the first to comment!

SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Related Skills

Caveman

Ultra-compressed communication mode. Cuts token usage ~75% by speaking like caveman while keeping full technical accuracy. Supports intensity levels: lite, full (default), ultra, wenyan-lite, wenyan-full, wenyan-ultra. Use when user says "caveman mode", "talk like caveman", "use caveman", "less tokens", "be brief", or invokes /caveman. Also auto-triggers when token efficiency is requested.

1023331 votes

Hyperplan

Adversarial multi-agent planning skill. Self-orchestrates 5 hostile category members (unspecified-low, unspecified-high, deep, ultrabrain, artistry) via team-mode for ruthless cross-critique debate, distills only the defensible insights, then MANDATORILY hands the distilled insight bundle to the `plan` agent for executable plan formalization. Use when planning needs maximum rigor and surfacing of weak assumptions, blind spots, and over-engineering. Triggers: 'hyperplan', 'hpp', '/hyperplan', ...

686011 votes

Mcp Code Execution

Routes multi-tool workflows through MCP servers for large datasets and pipelines. Use when Bash tool overhead is limiting throughput on data-heavy tasks.

3331 votes

catchup

Recovers prior coding-agent session context by running `catchup <agent> --since-compact`, which extracts a clean summary of a previous Codex, Claude Code, Antigravity, OpenCode, or Pi Agent session. Use when the user says "catch up", "what did the last session do", "get me up to speed", "I switched agents", or asks to recover/summarize a previous session before continuing. Do NOT use for the current conversation, git history, or any non-agent log.

611 votes

math-skill

A comprehensive mathematical reasoning skill for AI assistants — handles arithmetic to research-level problems with rigorous step-by-step reasoning, systematic verification, and transparent uncertainty handling

381 votes
View all in ai-agents →