Skills DirectorySkills Directory
SkillsLearnSecurityCategoriesDocsBlogPro
Sign InSubmit Skill
Skills Directory

Security-tested agent skills for Claude, coding agents, and AI workflows.

Directory

  • Browse Skills
  • All Skills A–Z
  • Claude Skills
  • Claude Code Skills
  • Agent Skills
  • Categories
  • Authors
  • Submit a Skill

Learn

  • Learn Hub
  • Install Claude Skills
  • Write SKILL.md
  • Skills vs MCP
  • Directories Compared

Security

  • Security
  • Methodology
  • Secure Claude Skills
  • Security Badges
  • Chrome Extension
  • Skill Manager

Company

  • About
  • Community
  • Blog
  • API Docs
  • Advertise

2026 Skills Directory. All rights reserved.

ProTermsPrivacyRefunds
Back to skills

Verification Rubric

ASecurity

Use when judging a task_gates entry whose verification_type is semantic or self_review -- reading the gate's evidence_shape as an explicit rubric, assessing the produced work against each stated criterion, and emitting a justified pass/fail verdict. Loaded by the verifier role for judgment-based gates.

4 stars
0 votes
0 copies
0 views
Added 9/23/2026
ai-agentsgo

Security Analysis

A100/100

Pro scans all 2 files and shows the line behind each finding

Scanned 9/24/2026

$npx -y skills add metraton/gaia --skill verification-rubric --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Verification Rubric?

Add the live security badge to your README — it updates automatically with every re-scan.

Security grade badge for Verification Rubric
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/metraton-verification-rubric/badge)](https://www.skillsdirectory.com/skills/metraton-verification-rubric)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
Files
SKILL.md
---
name: verification-rubric
description: Use when judging a task_gates entry whose verification_type is semantic or self_review -- reading the gate's evidence_shape as an explicit rubric, assessing the produced work against each stated criterion, and emitting a justified pass/fail verdict. Loaded by the verifier role for judgment-based gates.
---

# Verification Rubric

Verification-rubric is the judging discipline for a `task_gates` entry whose
`verification_type` is `semantic` or `self_review`: it reads the gate's
`evidence_shape` as an explicit rubric, assesses the produced work against
each stated criterion independently, and emits a structured, justified
pass/fail verdict -- never a hollow "looks fine." It is the judgment-based
counterpart to the deterministic-oracle mode, which re-runs a `command`/`code`
gate and diffs the result; there is no command to re-run here, only evidence
to weigh against criteria.

## Core principle

Two of the four `verification_type` values (`gaia.state.VALID_VERIFICATION_TYPES`,
`gaia/state/__init__.py`) route to this skill: **`semantic`** (`evidence_shape`
carries an externally authored rubric -- the contract-envelope's own
`requires_human` marker names the same discipline: "needs human/rubric
validation") and **`self_review`** (`evidence_shape` is the producer's own
statement of what it checked -- the envelope's `reviewed` field is the same
idea, and it is the honesty-rule floor from `agent-protocol`, not a shortcut
around it). Both are LLM-as-judge tasks, not code execution.

Three forces shape the judgment:

- **Explicit criteria beat holistic impression.** Read the rubric as a list of
  discrete, independently-checkable claims, not one paragraph to eyeball. A
  verdict built on an overall feeling collapses the moment someone asks "which
  part, exactly?"
- **Judge each criterion on its own -- resist position and verbosity bias.**
  Do not let a strong criterion compensate for a failing one (verdict
  inflation), do not let whichever criterion you read first anchor the whole
  read (position bias), and do not treat a longer or more polished artifact as
  automatically meeting more criteria (verbosity bias) -- effort and length
  are not what the rubric asks for.
- **Confirmed beats assumed.** A criterion is "met" only when the produced
  work was actually observed to satisfy it. An unobserved criterion is not
  met -- it is unknown, and unknown is not pass (mirrors `investigation`'s
  confirmed-vs-assumed line).

## The judging cycle

1. **Load the gate.** Read the `task_gates` row: `verification_type`
   (`semantic`|`self_review`), `evidence_shape` (the rubric prose or
   self-review statement), `artifact_path` (the produced work, if any). Run
   `gaia.state.gate_validation.validate_gate` first -- a structurally invalid
   gate (missing/empty `evidence_shape`) has nothing to judge; return that
   rejection, not a verdict. A gate with `stale_at` set keeps an old verdict
   that predates a change to the gate, its task or a covered AC: judge it
   afresh, as if pending.
2. **Parse the rubric into criteria.** Split `evidence_shape` into discrete,
   checkable statements -- one criterion, one claim. A rubric with a single
   paragraph is read as one criterion; a well-authored rubric names several,
   and each is graded independently.
3. **Gather the evidence.** Read the produced work at `artifact_path` (or
   wherever the task's outcome lives). Judge from what actually exists, never
   from the task's own description of itself.
4. **Assess each criterion independently.** For each: is it met, and why --
   the observation that grounds the call. An assessment with no stated reason
   is an assertion, not a judgment.
5. **Aggregate to the overall verdict.** Pass requires every criterion met,
   unless the rubric itself marks one optional/advisory (name that
   explicitly if so). One unmet, unwaived criterion is fail, regardless of how
   many others passed.
6. **Emit the structured, justified verdict** -- `{gate_id, verdict, criteria,
   overall_reasoning}` (see Reference implementation below). Justification is
   not optional: the rubric exists so the verdict can be defended
   criterion-by-criterion, not merely announced. A fail is recorded with the
   cause that says what has to move -- `product` (the work misses a
   criterion), `broken_test` (the rubric is uncheckable or contradicts
   itself), `requirement_changed` (the rubric no longer matches the brief);
   `environment` rarely applies to a judgment. The observations go in as
   evidence tied to the gate, `--negative` when they refute it.

## `self_review` vs `semantic`

- **`self_review`** -- `evidence_shape` is the producer's own statement of
  what it checked, not a rubric someone else authored. Judge it by the same
  discipline: does the statement name concrete checks and observations, or is
  it a hollow "looks good"? A statement that names nothing checked fails the
  honesty-rule floor regardless of how confident it reads.
- **`semantic`** -- `evidence_shape` is an externally authored, typically
  multi-criterion rubric -- the case this skill's cycle is built around.

## Reference implementation

`skills/verification-rubric/scripts/rubric_verdict.py` provides the deterministic assembly half of the
cycle: `skills/verification-rubric/scripts/rubric_verdict.py::parse_rubric_criteria` splits rubric prose into discrete criteria
(step 2), and `skills/verification-rubric/scripts/rubric_verdict.py::assemble_verdict` aggregates a list of `CriterionAssessment`
(`criterion`, `met`, `reasoning`) into a `RubricVerdict` (`verdict`,
`criteria`, `overall_reasoning`), rejecting an empty assessment list and any
assessment with blank `reasoning` -- the honesty rule enforced structurally,
mirroring `gaia.state.gate_validation.validate_gate`'s pure/deterministic
style. The per-criterion `met` call -- reading the evidence against the
criterion's wording -- is the judge's job (step 4); this module only
assembles what has already been judged, it does not automate the judgment.

## Anti-patterns

- **Holistic pass** -- grading the whole artifact by one overall impression
  instead of per-criterion. A strong majority does not offset one unmet
  criterion.
- **Hollow verdict** -- `pass`/`fail` with no stated reasoning. The rubric
  exists so the verdict can be defended, not merely declared.
- **Verbosity bias** -- crediting a longer or more detailed artifact as
  automatically meeting more criteria. Judge against the criterion's wording,
  not against effort.
- **Position bias** -- anchoring the whole judgment on whichever criterion was
  read first instead of giving each an independent look.
- **Treating `self_review`'s floor as license to skip observation** --
  `self_review` still requires naming what was checked; "I reviewed it, it's
  fine" is a shrug, not a review.

Attribution

metratonmetraton
View sourceSee grades on GitHubMore from metraton →
SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments (0)

No comments yet. Be the first to comment!

SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Related Skills

Caveman

Terse caveman voice: answer first, fluff gone, every technical fact kept. Use for /caveman, "caveman mode", "talk like caveman", "be brief", "less tokens". Stays on until "stop caveman" or "normal mode".

1100021 votes

Hyperplan

Adversarial multi-agent planning skill. Self-orchestrates 5 hostile category members (unspecified-low, unspecified-high, deep, ultrabrain, artistry) via team-mode for ruthless cross-critique debate, distills only the defensible insights, then MANDATORILY hands the distilled insight bundle to the `plan` agent for executable plan formalization. Use when planning needs maximum rigor and surfacing of weak assumptions, blind spots, and over-engineering. Triggers: 'hyperplan', 'hpp', '/hyperplan', ...

698461 votes

Writing Skills

Create and manage Claude Code skills in HASH repository following Anthropic best practices. Use when creating new skills, modifying skill-rules.json, understanding trigger patterns, working with hooks, debugging skill activation, or implementing progressive disclosure. Covers skill structure, YAML frontmatter, trigger types (keywords, intent patterns), UserPromptSubmit hook, and the 500-line rule. Includes validation and debugging with SKILL_DEBUG. Examples include rust-error-stack, cargo-dep...

3931 votes

Mcp Code Execution

Routes multi-tool workflows through MCP servers for large datasets and pipelines. Use when Bash tool overhead is limiting throughput on data-heavy tasks.

3421 votes

catchup

Recovers the conversation and failed tool calls of a previous Codex, Amp, Claude Code, Antigravity, Cline, Copilot CLI, Cursor, DeepSeek Harness, Grok Build, Kimi, OpenCode, Pi Agent, or ZCode session. Use when the user says "catch up", "what did the last session do", "get me up to speed", "I switched agents", asks to recover/summarize a previous session before continuing, or asks to diagnose or report a catchup failure. Do NOT use for the current conversation, git history, or any non-agent log.

741 votes
View all in ai-agents →