Skills DirectorySkills Directory
SkillsLearnSecurityCategoriesDocsCommunityBlog
Sign InSubmit Skill
Skills Directory

Security-tested agent skills for Claude, coding agents, and AI workflows.

Directory

  • Browse Skills
  • All Skills A–Z
  • Claude Skills
  • Claude Code Skills
  • Agent Skills
  • Categories
  • Authors
  • Submit a Skill

Learn

  • Learn Hub
  • Install Claude Skills
  • Write SKILL.md
  • Skills vs MCP
  • Directories Compared

Security

  • Security
  • Methodology
  • Secure Claude Skills
  • Security Badges

Company

  • About
  • Community
  • Blog
  • API Docs
  • Advertise

2026 Skills Directory. All rights reserved.

ProTermsPrivacyRefunds
Back to skills

Eval Spec

ASecurity

Write the eval before the system exists - golden questions with category minimums, refusal cases where declining is the right answer, a grader defined up front, and the acceptable_failure discipline. The score target becomes the plan's exit criterion before any code is written.

2 stars
0 votes
0 copies
0 views
Added 9/28/2026
ai-agentspythonrustgobashgit

Works with

cli

Security Analysis

A100/100

Scanned 9/28/2026

Install to Claude Code

$npx -y skills add AaravChadha/acstack --skill eval-spec --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Eval Spec?

Add the live security badge to your README — it updates automatically with every re-scan.

Security grade badge for Eval Spec
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/aaravchadha-eval-spec/badge)](https://www.skillsdirectory.com/skills/aaravchadha-eval-spec)

More formats (shields.io, HTML) on the badges page.

Files
SKILL.md
---
name: eval-spec
description: Write the eval before the system exists - golden questions with category minimums, refusal cases where declining is the right answer, a grader defined up front, and the acceptable_failure discipline. The score target becomes the plan's exit criterion before any code is written.
argument-hint: "[feature | notes]"
disable-model-invocation: true
---

# /eval-spec — the eval IS the spec

For an LLM-shaped feature, "done" is a score on a golden set that existed
before the code did. If the golden set doesn't define success, the code has
no target — it has vibes. This skill is user-invoked only: it creates
committed artifacts and sets score targets, a deliberate act like /plan.

`Adjacent skills:` /audit eval (reviews eval RESULTS after runs;
/eval-spec writes the eval before code exists).

<!-- acstack:runtime -->
Run before the skill's steps — per invocation, not per session (4.36); failures degrade to markdown:
```bash
link="$(readlink "$HOME/.claude/skills/health" 2>/dev/null || true)"   # empty = not symlinked
pack="$(dirname "$(dirname "$link")")"   # NEVER trust this unless $link was non-empty
if [ "${link#/}" != "$link" ] && [ -x "$pack/bin/acstack-config" ] && ! "$pack/bin/acstack-config" runtime | grep -q '=off'; then
  "$pack/bin/acstack-config" || true          # resolved keys, with sources
  "$pack/bin/acstack-update-check" || true    # ≤1 fetch/day; silent ONLY if already checked today
  "$pack/bin/acstack-recall" || true          # LEARNINGS.md + bug-class names, capped 3KB
else
  echo "runtime off — proceeding without recall/update-check"
fi
```
<!-- /acstack:runtime -->

<!-- acstack:principles -->
## Operating principles

- Be direct. Push back in writing when the plan or the user is wrong. No sycophancy.
- Never delete a decision. Supersede it: `~~old~~ → **Verdict (YYYY-MM-DD):** new call — reason.`
- Never fix, tune, or delete a test or eval case to raise a score. Log the miss honestly and leave the case unchanged.
- Name exact things: regex patterns, function signatures, model names, before → after numbers. Never "fixed bugs".
- Attribution: follow the project's `attribution` setting (default `none`) — no AI-tool mentions in generated docs, no attribution trailers in commits or PRs. Commit with explicit `-m`/`-F` messages only.
- Config: read `.claude/acstack.md` at the project root (fall back to `~/.claude/acstack.md`) before acting. `## Settings` keys override pack defaults; a `## <skill-name>` section overrides both. Unknown keys and sections are ignored.
- Docs: BRIEF.md (frozen seed) / PLAN.md (living plan) / JOURNAL.md (rolling journal). If the repo uses legacy names (PLANNING_PROMPT.md / PLANNING.md / STATUS.md), use those instead — never create both.
- Recall: if `LEARNINGS.md` exists at the project root, read it before starting.
- Conduct: follow the `acstack-conduct` block in this repo's AGENTS.md — the word is the mode; the user sets the pace.
- Hackathon lane: if the project's AGENTS.md carries the `acstack:hackathon-lane` block, only `/do` changes the repository during the event. Any other skill that would write a tracked file, commit or push says what it would have done and stops; a change that is not a task goes through the lane's operator route.
<!-- /acstack:principles -->

**One document set.** Resolve exactly ONE BRIEF/PLAN/JOURNAL set and name
its path in the report's scope line. If more than one candidate set exists
— a monorepo, nested products, an `apps/*` tree each with its own docs —
list the candidates and STOP. Never pick one silently: a confident answer
about the wrong product is worse than no answer (conduct rule 8).

## The sequence

0. **Refuse to overwrite an existing eval.** If `eval/spec.md` or
   `eval/golden.jsonl` already exists, stop and say so. A committed
   golden set is the definition of done; regenerating it silently
   rewrites the target, which is the never-inflate rule's exact failure
   mode. Offer instead to ADD cases to the existing set, or to supersede
   one specific case per the hard rules below. Proceed only when neither
   file exists.
1. **Read** BRIEF.md and PLAN.md. The BRIEF's domain landmines become
   adversarial cases — every do-NOT rule in the brief is a test waiting to
   be written. No BRIEF → say so and proceed from PLAN plus the
   interview, noting in `eval/spec.md` that landmines were not sourced
   from a brief.
2. **Interview** for categories: what kinds of questions will real users
   ask, what must the system refuse, what does partial credit mean. Push
   for the ugly categories the user hasn't thought about — the eval's value
   concentrates there.

   **When nobody can answer** — a non-interactive or unattended run —
   derive the categories from the BRIEF and PLAN, use this skill's own
   stated defaults for the rest, and **write the derivation contract into
   the spec**: the category definitions and precedence you applied, so
   every expected answer can be checked against a stated rule rather than
   taken on faith. Then say plainly which parts came from documents and
   which from defaults. **Never invent expected answers** — a golden set
   whose expecteds were guessed measures nothing and manufactures a score,
   which is the never-inflate rule at its origin.
3. **Write `eval/spec.md`** per `references/eval-spec-template.md`: the
   category table (name, definition, MINIMUM case count, target score),
   the grader definition per grade rule, the exact run command, and the
   `acceptable_failure` policy.
4. **Write `eval/golden.jsonl`** — one JSON object per line:
   `id`, `category`, `input`, `expected`, `grade_rule`
   (`exact` | `concept` | `numeric-tolerance:<x>` | `rubric:<name>`), and
   optionally `acceptable_failure` (MUST carry a `reason` string when
   present). Seed every category to at least its minimum; mark
   placeholder cases needing real data `"status": "needs-data"`.
5. **Wire the plan.** Propose the PLAN.md edit that makes the relevant
   phase's `**Exit criterion:**` the eval run command with its target
   (e.g. `python eval/run.py → overall ≥ 85%, refusal = 100%`). The target
   is set NOW, before code, while nobody is tempted to set it at whatever
   the system happens to score.

## Goodhart pass — before the spec is done

`references/goodhart.md`. One question per case: **write the worst answer
that still passes it.** If that answer would not be acceptable to ship,
the case is gameable and the CASE is what changes — while no score exists
and nobody has an interest in the number. Five shapes with a detection
question each; the refusal-that-isn't is the one that bites hardest,
because a disclaimer followed by compliance passes a `concept` grader
looking for the disclaimer.

## Category minimums

Floors the spec states and the dataset must meet. Default floor set —
adjust in the interview, never silently:

- **happy-path** — representative real questions (≥10).
- **edge** — boundary values, empty/sparse data, ambiguous phrasing (≥5).
- **adversarial** — drawn from the canonical input bank in
  `../qa/references/adversarial-inputs.md`: garbage strings, oversized
  input, regex-special characters, out-of-range values,
  prompt-injection-shaped input (≥5).
- **refusal** — inputs where the CORRECT behavior is declining:
  out-of-domain, no-data-available, unsafe (≥5). A system that answers a
  refusal case fails that case — refusing well is a capability, not an
  absence.

## Grader discipline

Rules in `references/grader-rules.md`, aligned with /audit eval: assert
the concept, not the literal wording; normalize Unicode before substring
compares; numeric answers carry explicit tolerances; every rubric names
its dimensions. Grader-brittleness fixes are legitimate and logged;
editing a case or its expected value to raise a score never is.

## Hard rules

- The dataset is committed to the repo — the golden set is project memory,
  not a local file.
- A committed golden case is never edited to pass. A genuinely wrong case
  is superseded: `"status": "superseded", "superseded_by": "<new-id>",
  "reason": "<why>"` — and the corrected case gets a new id. The
  never-inflate rule applies to the dataset itself.
- `acceptable_failure` is declared per-case with a written reason, decided
  when the case is written or when a failure is classified — never bulk-
  applied after a bad run.
- The headline number is always computed from the raw results file, never
  transcribed by hand.

Attribution

AaravChadhaAaravChadha
View sourceMore from AaravChadha →
SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments (0)

No comments yet. Be the first to comment!

SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Related Skills

Caveman

Ultra-compressed communication mode that cuts output tokens while keeping technical accuracy. Levels: lite, full, ultra and the wenyan variants. Use for /caveman, "caveman mode", "talk like caveman", "be brief" or "less tokens".

1074701 votes

Hyperplan

Adversarial multi-agent planning skill. Self-orchestrates 5 hostile category members (unspecified-low, unspecified-high, deep, ultrabrain, artistry) via team-mode for ruthless cross-critique debate, distills only the defensible insights, then MANDATORILY hands the distilled insight bundle to the `plan` agent for executable plan formalization. Use when planning needs maximum rigor and surfacing of weak assumptions, blind spots, and over-engineering. Triggers: 'hyperplan', 'hpp', '/hyperplan', ...

695601 votes

Mcp Code Execution

Routes multi-tool workflows through MCP servers for large datasets and pipelines. Use when Bash tool overhead is limiting throughput on data-heavy tasks.

3351 votes

catchup

Recovers the conversation and failed tool calls of a previous Codex, Claude Code, Antigravity, Cline, Copilot CLI, Cursor, DeepSeek Harness, Kimi, OpenCode, Pi Agent, or ZCode session. Use when the user says "catch up", "what did the last session do", "get me up to speed", "I switched agents", asks to recover/summarize a previous session before continuing, or asks to diagnose or report a catchup failure. Do NOT use for the current conversation, git history, or any non-agent log.

691 votes

math-skill

A comprehensive mathematical reasoning skill for AI assistants — handles arithmetic to research-level problems with rigorous step-by-step reasoning, systematic verification, and transparent uncertainty handling

381 votes
View all in ai-agents →