Skills DirectorySkills Directory
SkillsLearnSecurityCategoriesDocsBlogPro
Sign InSubmit Skill
Skills Directory

Security-tested agent skills for Claude, coding agents, and AI workflows.

Directory

  • Browse Skills
  • All Skills A–Z
  • Claude Skills
  • Claude Code Skills
  • Agent Skills
  • Categories
  • Authors
  • Submit a Skill

Learn

  • Learn Hub
  • Install Claude Skills
  • Write SKILL.md
  • Skills vs MCP
  • Directories Compared

Security

  • Security
  • Methodology
  • Secure Claude Skills
  • Security Badges
  • Chrome Extension
  • Skill Manager

Company

  • About
  • Community
  • Blog
  • API Docs
  • Advertise

2026 Skills Directory. All rights reserved.

ProTermsPrivacyRefunds
Back to skills

Evaluator

ASecurity

Externally evaluates an already-implemented feature against its contract.md (environment, quality gates, coverage manifest, observable criteria), producing screenshots, a report, and chat findings. Owns the evaluation loop — keeps the attempt counter in progress.json, decides the state (CLEAN/FAIL/PENDING/ABORTED), and routes each failure: code failures (gate/observable) to fix-runner, test failures to the matching test-writer (then confirmed by the matching test-validator) before resuming. L...

2 stars
0 votes
0 copies
1 views
Added 9/19/2026
ai-agentsgogit

Works with

cli

Security Analysis

A100/100

Pro scans all 2 files and shows the line behind each finding

Scanned 9/19/2026

$npx -y skills add dayvisonassis/sdd-skills --skill evaluator --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Evaluator?

Add the live security badge to your README — it updates automatically with every re-scan.

Security grade badge for Evaluator
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/dayvisonassis-evaluator/badge)](https://www.skillsdirectory.com/skills/dayvisonassis-evaluator)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
Files
SKILL.md
---
name: evaluator
description: Externally evaluates an already-implemented feature against its contract.md (environment, quality gates, coverage manifest, observable criteria), producing screenshots, a report, and chat findings. Owns the evaluation loop — keeps the attempt counter in progress.json, decides the state (CLEAN/FAIL/PENDING/ABORTED), and routes each failure: code failures (gate/observable) to fix-runner, test failures to the matching test-writer (then confirmed by the matching test-validator) before resuming. Loops until CLEAN, PENDING, or ABORTED.
---

# Evaluator

A second, independent layer of validation. After `implement-feature` finishes a feature, the `evaluator` looks at it **from the outside** and confirms — against `contract.md` — whether the delivery conforms. It finds deviations the implementer's self-check missed. It **does not fix code**; it evaluates and **orchestrates the correction loop**, dispatching `fix-runner` on failure.

> Base docs: `https://github.com/dayvisonassis/sdd-skills/blob/main/docs/Skill_Evaluator.md` (full rationale),
> `https://github.com/dayvisonassis/sdd-skills/blob/main/docs/Contrato_de_Feature.md` (contract structure),
> `https://github.com/dayvisonassis/sdd-skills/blob/main/docs/Como_criar_gates.md` (gates),
> `https://github.com/dayvisonassis/sdd-skills/blob/main/docs/Fluxo_SDD_e_Implementacao_das_Skills.md` (flow, states, progress.json schema).
> Report schema: `references/evaluation-report-schema.md`.

The evaluator complements automated tests — it does not replace them.

## INPUT

Free-form. The skill needs:

- **The target feature** (ID/name) — explicit and required. Abort if absent or ambiguous (list candidates).
- Auto-discovers: `progress.json` (root of `docs/`) and the feature's `contract.md` (in `docs/<feature-id>-<kebab>/`).
- Optional free-form overrides at the end — e.g. "no screenshots", "gates only", "max 5 attempts" (overrides `maxFixAttempts` for this run).

If the feature has no `contract.md`, abort: "No contract.md for F<ID> — generate it with `spec-writer` first."

## OUTPUT

- **Screenshots** — visual evidence of observable UI criteria (when applicable).
- **Report** — consolidated ✓ / ✗ / — per contract criterion and gate.
- **Findings in chat** — textual summary for the user.
- **State in `progress.json`** — CLEAN | FAIL | PENDING | ABORTED for the feature.
- **On FAIL:** a structured `evaluation-report.json` in the feature folder, with each failure classified by `kind` (consumed by `fix-runner` for code, or a test-writer for tests).
- The evaluator does **not** edit code or tests itself. Code corrections → `fix-runner`; test corrections → the matching test-writer (confirmed by the matching test-validator).

---

## EXECUTION STEPS

### Step 1: Resolve Input

- Resolve the target feature (ID/name). Abort if absent or ambiguous, listing candidates.
- Locate `progress.json` (root of `docs/`) and the feature's `contract.md`. If `contract.md` is missing → abort ("generate the contract first with `spec-writer`"). If `progress.json` is missing, create it with `config.maxFixAttempts` default 3.
- Parse any overrides (e.g. `max N attempts`, `gates only`, `no screenshots`) and record them for the report.

### Step 2: Load Context

- Read `contract.md`: Environment Contract, Quality Gates (with `id`s), Coverage Manifest, Surfaces & Behaviors, Observable Criteria (with `id`s).
- Read `progress.json`: this feature's `state`, `attempt`, and `config.maxFixAttempts` (N).
- The contract is the **single source of acceptance criteria**. Do NOT invent criteria outside it.

### Step 3: Verify Environment Contract

- Check each environment prerequisite (runtime up, services reachable, tools available — e.g. dev server serving `/`, Playwright CLI present).
- If the environment is **not** satisfied → do NOT proceed to evaluation. Set `state: PENDING` (human/environment intervention needed), explain that this is an **environment** problem (not an implementation failure), and stop. Environment-invalid ≠ implementation-wrong.

### Step 4: Run Quality Gates

- Execute each gate declared in the contract (typecheck, lint, build, tests, arch...). Use the exact commands from the contract.
- Any gate failing is a contract violation. Collect the command + log as evidence.
- **Classify each failure by cause** (this drives the correction routing in Step 7):
  - The failure is a **code failure** (`kind: gate` or `observable-criterion`) when production code is wrong — a type error, lint/build/arch violation, or a missing observable behavior.
  - The failure is a **test failure** (`kind: test`) when the test itself is broken/non-conforming — e.g. the `tests` gate fails and the cause is the test file (missing/incorrect mock, wrong pattern, a test that no longer matches correct behavior), not the production code. For a `kind: test` failure, fill the routing fields (`testSuite`, `testFile`, `targetFile`) per `references/evaluation-report-schema.md`.
  - When ambiguous (a failing test that might reflect a real code bug), prefer `kind: gate`/`observable-criterion` and let `fix-runner` handle the code; only route to a test-writer when the test is clearly the thing that is wrong.
- (If `gates only` override is set, skip Step 5 and go to Step 6 with just gate results.)

### Step 5: Validate Surfaces & Observable Criteria

- For each surface in the Coverage Manifest, start from its declared initial state and exercise the concrete behaviors (e.g. navigate the route via Playwright CLI, capturing screenshots).
- For each Observable Criterion, collect verifiable evidence (e.g. CTA present, login in top nav, redirect behavior, visual identity). A criterion with no observable evidence is a failure (`kind: observable-criterion`, `ref: <crit-id>`).
- Map every PRD-derived acceptance back to a contract criterion/gate (the contract already did this traceability; honor it).

### Step 6: Decide State

Determine the feature's state strictly from contract adherence — never from "looks ok":

- **CLEAN** — no failures: all gates pass, all observable criteria met. → record state, proceed to Step 9.
- **FAIL** — at least one **correctable** failure (gate/test/observable). → go to Step 7 (loop).
- **PENDING** — something the evaluator **cannot test by itself** (needs a human, or an environment it cannot bring up). → record state with a note, proceed to Step 9.
- **ABORTED** — decided in Step 7 when attempts are exhausted.

List exactly which gates/criteria failed.

### Step 7: Correction Loop (when FAIL) — route by failure kind

- Read `attempt` and `maxFixAttempts` (N) from `progress.json`.
- **If `attempt >= N`** → set `state: ABORTED`; stop and report (the loop tried N times without converging). Proceed to Step 9.
- **Else:**
  1. Write/refresh `evaluation-report.json` in the feature folder (schema in `references/evaluation-report-schema.md`) with the current `attempt` and the classified `failures[]`.
  2. Set `state: FAIL` in `progress.json` and persist the report path in `lastEvaluationReport`.
  3. **Route each failure by `kind`:**
     - **`kind: gate` / `observable-criterion` (code)** → **dispatch `fix-runner`**, passing the feature ID and the report path. Unchanged behavior — the on-disk report is the source of truth.
     - **`kind: test`** → run the **test-correction sub-flow (Step 8)** for that failure. Never send test failures to `fix-runner`.
  4. When the dispatched correction returns:
     - Correction applied → **increment `attempt`** in `progress.json`, then **re-evaluate**: go back to Step 3.
     - "not resolved — <reason>" → still increment `attempt`; if `attempt >= N` set `ABORTED`, else re-evaluate. Do not loop without incrementing.

The evaluator **owns** the counter, the limit N, and the ABORTED decision. Correction skills are stateless and never decide when to stop.

### Step 8: Test-correction sub-flow (`kind: test`)

For a test failure, correcting the code is the wrong move — fix the test, then re-confirm it
conforms. Select the suite from `testSuite`/`testFile` (deterministic rule in the schema):
`unit` → `unit-test-*`, `integration` → `integration-test-*`, `monorepo` → `monorepo-unit-test-*`.

1. **Fix the test** — dispatch the matching **test-writer** in **correction mode** (autonomous),
   passing the feature ID, the `evaluation-report.json` path, `testFile`, and `targetFile`. It
   fixes only the flagged test (smallest footprint), never production code.
2. **Confirm conformance** — dispatch the matching **test-validator** on the corrected `testFile`.
   - Verdict **PASS** (or PASS WITH WARNINGS) → the test now conforms; **resume the evaluation
     where it left off** (re-run Step 3+ / the failing gate) and continue.
   - Verdict **FAIL** → the correction did not conform. Treat this round as a spent attempt:
     increment `attempt`; if `attempt >= N` → `ABORTED`; else loop (dispatch the test-writer again
     with the validator's findings, then re-validate).
3. Only after the test-validator returns PASS does the evaluator continue its own evaluation.

> The test-writer/validator pair is a **sub-loop inside** the evaluator's main loop. It still
> consumes the single `attempt`/N budget — never iterate the sub-loop without incrementing.

### Step 9: Persist State & Report

- Write the final `state` (CLEAN | PENDING | ABORTED) and `updatedAt` for the feature in `progress.json`. Preserve `config` and other features' entries (merge, don't replace).
- On CLEAN, you may clear or keep `lastEvaluationReport` (a stale FAIL report should not imply a current failure — prefer clearing it).
- Output the report to chat:

```
Feature F<ID> — <name>

State: CLEAN | FAIL→(looped) | PENDING | ABORTED
Attempts used: <attempt> / <maxFixAttempts>

Quality gates:
✓ <gate-id> passed
✗ <gate-id> failed: <message> (<command>)

Observable criteria:
✓ <crit-id> — <evidence / screenshot path>
✗ <crit-id> — <what was missing>

Surfaces evaluated:
- <surface-id>: <result>

Environment:
✓ contract met  |  ✗ not met → PENDING (<which prerequisite>)

Findings:
- <key deviations the evaluation found>

Next:
- CLEAN → feature ready for the next stage
- PENDING → needs human intervention: <what>
- ABORTED → tried <N> fixes without converging; see last evaluation-report.json
```

---

## RULES

**Always:**
- Treat `contract.md` as the single source of acceptance criteria.
- Abort the evaluation if the Environment Contract is not met, and say it is an environment problem (→ PENDING).
- Run the contract's Quality Gates as objective checks before/with functional inspection.
- Derive the state from contract adherence, with evidence per criterion.
- Own the loop: keep `attempt`/`maxFixAttempts` in `progress.json` and decide CLEAN/FAIL/PENDING/ABORTED.
- Classify each failure by kind and route it: `gate`/`observable-criterion` → `fix-runner`; `test` → the matching test-writer (then confirmed by the matching test-validator).
- After a `kind: test` correction, require the **test-validator PASS** before resuming the evaluation.
- Re-evaluate after each correction until CLEAN, PENDING, or ABORTED.
- Increment `attempt` once per correction round (code or test); never loop without incrementing.

**Never:**
- Alter production code or "fix" the feature yourself — code corrections are the `fix-runner`'s job, test corrections are the test-writers' job.
- Send a `kind: test` failure to `fix-runner`, or a `kind: gate`/`observable-criterion` failure to a test-writer.
- Approve based on visual perception without running the objective gates.
- Evaluate a feature without a `contract.md`.
- Invent criteria not present in the contract.
- Exceed `maxFixAttempts` without marking ABORTED (the test sub-loop shares the same budget).
- Write `PENDING_EVALUATION` into `progress.json` — that is the implementer's pre-state; the evaluator writes CLEAN/FAIL/PENDING/ABORTED.
- Confuse an environment failure (PENDING) with an implementation failure (FAIL).

---

## Edge Cases

**No contract.md for the feature**: abort and direct the user to `spec-writer`.

**progress.json missing**: create it with `config.maxFixAttempts` = 3 and this feature's entry.

**Environment Contract not met**: state = PENDING, clearly labeled as environment, no fix-runner dispatch.

**Gate command not runnable in this environment** (e.g. missing toolchain): treat as PENDING for that gate (needs environment), not FAIL — do not send the fix-runner after an environment gap. Note it in the report.

**maxFixAttempts reached**: state = ABORTED; keep the last `evaluation-report.json` for inspection.

**fix-runner reports "not resolved"**: increment `attempt`; abort to ABORTED if the limit is hit, otherwise re-evaluate.

**Override `max N attempts`**: use N for `maxFixAttempts` this run (and persist it to `config` if the user intends it to stick — otherwise apply for the session only and note it).

**Observable criterion is runtime-only and the runtime can't be exercised**: PENDING for that criterion (human/environment), not FAIL.

**Feature already CLEAN in progress.json**: re-running is allowed (idempotent re-check); report the result and refresh state.

**Cross-feature criteria**: if the contract references behavior provided by another feature, evaluate only this feature's surface; note cross-feature dependencies in findings rather than failing on another feature's gap.

Attribution

dayvisonassisdayvisonassis
View sourceSee grades on GitHubMore from dayvisonassis →
SSkills Directory ProSkills Directory

Get any skill into Claude in one click.

Download any skill as a ZIP for Claude.ai, Claude Desktop, or .claude/skills. $9/mo.

See Pro

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments (0)

No comments yet. Be the first to comment!

SSkills Directory ProSkills Directory

Get any skill into Claude in one click.

Download any skill as a ZIP for Claude.ai, Claude Desktop, or .claude/skills. $9/mo.

See Pro

Related Skills

Caveman

Ultra-compressed communication mode that cuts output tokens while keeping technical accuracy. Levels: lite, full, ultra and the wenyan variants. Use for /caveman, "caveman mode", "talk like caveman", "be brief" or "less tokens".

1074701 votes

Hyperplan

Adversarial multi-agent planning skill. Self-orchestrates 5 hostile category members (unspecified-low, unspecified-high, deep, ultrabrain, artistry) via team-mode for ruthless cross-critique debate, distills only the defensible insights, then MANDATORILY hands the distilled insight bundle to the `plan` agent for executable plan formalization. Use when planning needs maximum rigor and surfacing of weak assumptions, blind spots, and over-engineering. Triggers: 'hyperplan', 'hpp', '/hyperplan', ...

696561 votes

Writing Skills

Create and manage Claude Code skills in HASH repository following Anthropic best practices. Use when creating new skills, modifying skill-rules.json, understanding trigger patterns, working with hooks, debugging skill activation, or implementing progressive disclosure. Covers skill structure, YAML frontmatter, trigger types (keywords, intent patterns), UserPromptSubmit hook, and the 500-line rule. Includes validation and debugging with SKILL_DEBUG. Examples include rust-error-stack, cargo-dep...

3931 votes

Mcp Code Execution

Routes multi-tool workflows through MCP servers for large datasets and pipelines. Use when Bash tool overhead is limiting throughput on data-heavy tasks.

3351 votes

catchup

Recovers the conversation and failed tool calls of a previous Codex, Claude Code, Antigravity, Cline, Copilot CLI, Cursor, DeepSeek Harness, Kimi, OpenCode, Pi Agent, or ZCode session. Use when the user says "catch up", "what did the last session do", "get me up to speed", "I switched agents", asks to recover/summarize a previous session before continuing, or asks to diagnose or report a catchup failure. Do NOT use for the current conversation, git history, or any non-agent log.

691 votes
View all in ai-agents →