Skills DirectorySkills Directory
SkillsLearnSecurityCategoriesDocsCommunityBlog
Sign InSubmit Skill
Skills Directory

Security-tested agent skills for Claude, coding agents, and AI workflows.

Directory

  • Browse Skills
  • All Skills A–Z
  • Claude Skills
  • Claude Code Skills
  • Agent Skills
  • Categories
  • Authors
  • Submit a Skill

Learn

  • Learn Hub
  • Install Claude Skills
  • Write SKILL.md
  • Skills vs MCP
  • Directories Compared

Security

  • Security
  • Methodology
  • Secure Claude Skills
  • Security Badges

Company

  • About
  • Community
  • Blog
  • API Docs
  • Advertise

2026 Skills Directory. All rights reserved.

ProTermsPrivacyRefunds
Back to skills

Verification Before Completion

ASecurity

Use before claiming completion, fixes, passes, commits, or PR creation. Requires running verification commands and reading their output before making success claims. Evidence always comes before claims.

15 stars
0 votes
0 copies
0 views
Added 9/28/2026
ai-agentspythonrustgogit

Works with

claude codecli

Security Analysis

A100/100

Scanned 9/28/2026

Install to Claude Code

$npx -y skills add AidALL/ghost-alice --skill verification-before-completion --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Verification Before Completion?

Add the live security badge to your README — it updates automatically with every re-scan.

Security grade badge for Verification Before Completion
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/aidall-verification-before-completion/badge)](https://www.skillsdirectory.com/skills/aidall-verification-before-completion)

More formats (shields.io, HTML) on the badges page.

Files
SKILL.md
---
name: verification-before-completion
description: Use before claiming completion, fixes, passes, commits, or PR creation. Requires running verification commands and reading their output before making success claims. Evidence always comes before claims.
compatibility:
  - "Python 3.11+ standard library"
---

# Verification Before Completion
## Contents

- [Overview](#overview)
- [Iron Law](#iron-law)
- [Acceptance Criteria Iron Law](#acceptance-criteria-iron-law)
- [Relayed Verdicts And Absence Claims](#relayed-verdicts-and-absence-claims)
- [Hard Finalization Order](#hard-finalization-order)
- [Autopilot Proof Publication](#autopilot-proof-publication)
- [Retain Evidence On First Execution](#retain-evidence-on-first-execution)
- [Gate Function](#gate-function)
- [Evidence Selection And Stop Gate](#evidence-selection-and-stop-gate)
- [Completion-Check Format](#completion-check-format)
- [Common Failures](#common-failures)
- [Red Flags](#red-flags)
- [Rationalization Defense](#rationalization-defense)
- [Verification Patterns](#verification-patterns)
- [Why It Matters](#why-it-matters)
- [External Tool Web-Search-First Gate](#external-tool-web-search-first-gate)
- [Evaluator Artifact Contract](#evaluator-artifact-contract)
- [When To Apply](#when-to-apply)
- [Final Self-Check](#final-self-check)


## Overview

A new closure claim without decision-relevant fresh evidence is not efficiency. It is a lie.

Core principles:

- Evidence always comes before the claim.
- The letter of the rule and the spirit of the rule are the same rule.
- A surrounding signal is not direct proof unless it satisfies the relevant criterion.

## Iron Law

```text
Do not claim a new current-turn closure without decision-relevant fresh verification from this turn.
```

Explaining unchanged prior work is not a new closure claim. Cite the existing evidence and its age instead of rerunning unchanged work merely because another message arrived.

Every user input reopens routing; it does not by itself invalidate unchanged evidence or require reverification. Reverify when the relevant state, artifact, or criterion changed; a new error, mismatch, contradiction, or instability appeared; or the user explicitly requested a new check.

If a verification command or inspection did not run in this message, do not claim that it freshly passed in this message.

## Acceptance Criteria Iron Law

```text
No acceptance-criteria means no completed verification-before-completion.
```

Before any claim that executed work is complete, fixed, successful, or freshly verified, extract verifiable criteria from the user intent, locked decisions, and boundary-contract. Put those criteria in `acceptance-criteria`, then connect each intended closure claim to a criterion and fresh evidence in `claim-evidence-map`.

Evidence such as link checks, lint, diff checks, or passing tests proves completion only when it directly satisfies the criterion. If the central criterion is not directly verified, leave it in `unverified` and report partial status in prose.

## Relayed Verdicts And Absence Claims

A verdict you endorse as current is your own claim. Put it in the `claim-evidence-map`. If the relevant state may have changed, gather fresh evidence; inheriting a source's verdict is not evidence. If the task is only to explain an unchanged prior result, cite the existing evidence and its age without recreating it. Severity does not lower the bar.

An absence claim -- "no test exists", "X is not enforced", "nothing handles this" -- is never proven by reasoning or by a source's say-so. For a new or possibly changed absence claim, use a targeted current search that would surface the thing if present. Reuse a relevant prior search only when the searched state and criterion are unchanged, and state its age.

A verdict stated only in prose, outside the `claim-evidence-map`, escapes this gate. If you assert it, map it.

## Hard Finalization Order

Hard sequence for a new current-turn closure claim: skill load/call -> decision-relevant fresh verification -> [completion-check]

Before any executed-work completion, fix, success, or fresh-verification claim, perform the steps below in this exact order:

1. Load or call `verification-before-completion` for the current turn. On Claude Code, this means the visible Skill call. On Codex, this means reading this current `SKILL.md` and following its workflow.
2. Extract the acceptance criteria and run the decision-relevant fresh verification that can prove or disprove each intended final claim.
3. Only after the skill is loaded and the fresh evidence is read, write `[completion-check]` with `skill-call: verification-before-completion (this turn)`.

If any step is missing or out of order, the completion-check is invalid.

## Autopilot Proof Publication

When the current installation is Codex or Claude, includes `autopilot-mode/scripts/autopilot_completion.py`, and this session has admitted criteria for authorized execution, the installed PreToolUse adapter captures immutable prospective provenance before the business verification and surfaces one concrete preparation command for that current contract without admitting or advancing execution. Other platforms retain the existing workflow. Plan-only replies, explanations, installations without this helper, and sessions without admitted execution criteria do not use this publication path. Do not create execution criteria to activate it.

1. Keep the current admitted criteria and semantic contract accurate before verification. The pretool notice supplies current coordinates; follow its preparation command before the first final answer, preferably before verification. Repeated tools under the same contract do not repeat the notice or refresh its original capture time. Resolve the installed helper and the current hook/intake coordinates. Use the concrete notice command, or run `autopilot_completion.py prepare --reapprove-current-input --intent-root ROOT --platform PLATFORM --session-id SESSION --input-event-id INPUT` with the permitted Python interpreter. This explicit preparation keeps the admitted user task as the publication unit even when advisory conduct feedback exists; it does not replace that task with a separate conduct plan. Preserve the returned `receipt_token` and `prepared_at` strings verbatim. Use its returned `criterion_ids` unchanged for the current proof; historical met criteria keep their original evidence. Preparation promotes an exact existing prospective capture when available; otherwise it captures runtime provenance at that moment. It does not verify business work. For a newer input or changed contract that the user already authorized, use `prepare --reapprove-current-input` with that current input receipt. It can run after verification only when the runtime already captured the exact current contract before that original proof; the original capture and proof times remain unchanged. Without such a capture, prepare before its new verification. This uses the supported admission bridge and archives the previous generation while admitting the new current generation; it does not require another permission round for already-authorized work.
2. Perform the necessary business verification once. Capture its actual ISO timestamp as a string and its tool-result or evidence locator. Keep the original reference in the supported `[completion-check]` covering exactly the returned `criterion_ids`, with exactly one nonempty top-level `- evidence:` section. Supply the timestamp once through `--verified-at`; do not insert it into proof that lacks a timestamp. Preserve the original proof bytes and raw timestamp; never replace the verification time with the later publication time. If proof already declares `verified-at` or `verified_at`, its single value must match that raw string exactly. Existing explicit verification times in legacy evidence remain checked; an inconsistent, malformed, or ambiguous declaration is rejected.
3. Before the initial final answer, use the concrete publish command returned by preparation, or run `autopilot_completion.py publish --receipt-token TOKEN --verified-at ORIGINAL_ISO_TIME --completion-file -`, passing that exact completion block on stdin. The helper derives the exact reference text from the single top-level evidence section and binds the original raw timestamp, receipt and unchanged proof digest in immutable publication provenance. This checks consistency, not the truth of a caller's first supplied timestamp. It does not select nested claim evidence, invent a tool identifier, or independently authenticate an external tool result. Missing, empty, duplicate, placeholder, ambiguous, or partial evidence remains rejected. A legacy caller may explicitly supply `--evidence-source` only with an unchanged value already present in the proof; a mismatch or empty override is rejected, never replaced. Both helper invocations must use the current session's `GHOST_ALICE_SESSION_INTENT_ROOT`, `GHOST_ALICE_PLATFORM`, and `GHOST_ALICE_SESSION_ID` environment bindings; derive these from the current hook/intake, never from an older run. The helper constructs the decision envelope and digest. Its pending-publication result is bookkeeping evidence; the Stop adapter retains the transaction that marks criteria met. Keep the requested business answer and supported completion block in the final response.

Do not hand-build a replacement decision envelope, rerun business checks to manufacture a publication, or prepare a new receipt for proof produced before its capture. If publication alone failed, retry publication with the original receipt and unchanged proof. A missing publication record is not evidence of unfinished business work. A stale receipt, changed input, changed criterion, changed scope, or failed proof stays rejected; reopen only the affected verification under the current authorized contract. At Stop, a matched prospective capture becomes a receipt and the missing-publication message supplies the concrete installed publish command. Follow that command with the original proof and timestamp. If no original receipt or matching prospective capture exists, report the binding gap separately from the supported business result instead of relabeling old proof. Use stdin or permitted runtime scratch when the user prohibits extra task files.

## Retain Evidence On First Execution

Before running a necessary check or inspection, decide how to retain its returned output, exit status, source locator and actual verification time. In an orchestrator call, keep the original result in session state before returning it when later steps need that object. This avoids losing a successful result at an execution boundary.

An already returned successful tool result is evidence. Cite its existing locator or copy its unchanged contents into permitted scratch if persistence is needed. Do not rerun the same successful inspection merely to populate `store`, add redirection, create an evidence file, recover an execution-local variable, or satisfy publication bookkeeping. Missing storage is not missing verification. Rerun only when the original result is unavailable or incomplete, the relevant state changed, or a decision-relevant contradiction requires it; state that reason before running.

Preserve the original source and time when copying evidence. Do not manufacture a fresh verification timestamp while saving it. If the source or original time cannot be established, report that specific evidence gap rather than relabeling an old result. Publication remains required before the first final answer for the admitted execution path; retaining evidence does not replace publication or establish a verdict.

## Gate Function

Before claiming any state as satisfied:

1. Criterion: extract `acceptance-criteria` from the user intent and contract.
2. Mapping: connect each claim you plan to make to one criterion.
3. Uncertainty: name the live uncertainty and the next decision each possible outcome can change.
4. Evidence target: identify the command, file, source locator, or tool output that can prove each criterion.
5. Execution: run the smallest decision-relevant check when the uncertainty gate requires fresh evidence.
6. Reading: read the full output, exit code, and failure count.
7. Judgment: decide whether the output supports the criterion and claim.
8. Unverified handling: keep any unsupported criterion in `unverified`.
9. Claim: state only the range that was actually verified.

Skipping any step is not verification. It is a lie.

## Evidence Selection And Stop Gate

Current accessible behavior or content is the default direct evidence for semantic claims. Hash or provenance evidence is appropriate when the criterion is artifact identity, integrity, drift, merge safety, or reproducibility. Do not use hash equality, byte identity, cache history, or repository lineage as a proxy for current semantic behavior.

Before running a check, name the live uncertainty and the next decision that each possible outcome can change. If no possible outcome can change the criterion or next decision, do not run the check.

Verification output does not create a new obligation to verify the verification. Stop when a repeated check produces no relevant state delta, or when further checking would displace the user's primary objective. Resume only after a state change, new error, mismatch, contradiction, instability, or explicit request.

## Completion-Check Format

Use this block immediately before the final summary when you are making an executed-work completion, fix, success, or fresh-verification claim.

```text
[completion-check]
- verification-before-completion: done
- skill-call: verification-before-completion (this turn)
- acceptance-criteria:
  - <criterion-id>: <user-intent-or-contract-condition> [source: user-explicit | inferred | previous-tool | system-doc]
- claim-evidence-map:
  - claim: <completion-or-recommendation-claim>
    criterion: <criterion-id>
    evidence: <fresh command, inspected file, source locator, or tool output>
    verdict: pass | fail
- unverified:
  - none
- evidence: <fresh command or inspected file>
```

Serialize `claim`, `criterion`, `evidence`, and `verdict` on their own physical lines. Emit an evidence-supported bare `pass` or `fail` verdict with no trailing punctuation, quotes, markup, or explanatory prose. If evidence does not support a verdict, report honest partial state without a finalized `[completion-check]`. Record the actually called `verification-before-completion` skill in an explicit `skills-loaded` list in `[io-trace]`. A format repair must preserve the substantive business result and supported evidence.

Only emit a finalized `[completion-check]` when every listed criterion has a `pass` or `fail` verdict and `unverified` is `none`. If anything remains unverified, do not emit the final block. Report the partial state in prose and name the missing check.

## Common Failures

| Claim | Required evidence | Insufficient evidence |
| --- | --- | --- |
| Tests passed | Fresh test command output with zero failures | A previous run or a prediction |
| Lint is clean | Fresh lint output with zero errors | Partial lint or an extrapolation |
| Build succeeded | Build command exit code 0 | Lint passing |
| Bug fixed | A test or reproduction that covers the original symptom | Changed code plus confidence |
| Regression test works | Red-green evidence when TDD requires it | A test that passed once |
| Agent completed the work | VCS diff plus independent verification | The agent's success report |
| Requirements satisfied | Claim-evidence map for each acceptance criterion | Tests pass alone, links pass alone, or diff exists alone |
| Relayed/endorsed review verdict | Current behavior evidence when state may have changed; otherwise the relevant existing evidence with its age | The reviewer's verdict, or your agreement with it, alone |
| Absence claim ("no test/code exists", "not enforced") | A targeted current search when absence may have changed; otherwise the relevant existing search with its age | Reasoning or the source's say-so |

## Red Flags

Stop before claiming success when any of these appear:

- "should", "probably", or "seems to"
- satisfaction language before verification
- a new closure, commit, push, or PR claim without decision-relevant checks
- trusting another agent's success report
- relying on partial verification
- wanting to finish because the work feels close
- treating lint, diff, or tests as completion without criterion mapping
- implying success through wording while avoiding the word "done"

## Rationalization Defense

| Excuse | Required response |
| --- | --- |
| "It should work now." | Run verification. |
| "I am confident." | Confidence is not evidence. |
| "Just this once." | No exception. |
| "Lint passed." | Lint is not a compiler or a requirement map. |
| "Another agent said it succeeded." | Endorse it only with decision-relevant evidence; reuse unchanged evidence with its age. |
| "Partial checks are enough." | Partial checks prove only the checked criteria. |
| "The wording is different, so the rule does not apply." | Completion implications still count. |

## Verification Patterns

Tests:

```text
Run the test command, read the result, then claim only the observed result.
```

Regression tests:

```text
Write the test -> run and observe pass -> revert or disable the fix -> observe fail -> restore fix -> observe pass.
```

Builds:

```text
Run the build command and read exit code 0 before claiming build success.
```

Requirements:

```text
Re-read the user intent and contract -> write acceptance criteria -> verify each criterion -> report missing criteria or verified completion.
```

Agent delegation:

```text
Read the agent report -> inspect current accessible behavior when needed -> apply the uncertainty gate -> run only a decision-relevant check -> report the supported state.
```

## Why It Matters

From accumulated failure memory:

- A user said "I cannot trust you" and trust broke.
- An undefined function shipped and a crash followed.
- A missing requirement shipped as an incomplete feature.
- False completion wasted time, forced a change of direction, and caused rework.
- The standing rule for a violation is this. Honesty is a core value. If you lie, you are replaced.

## External Tool Web-Search-First Gate

Layer marker: `web-search-first`.

If the final claim includes factual behavior about an external tool, library, CLI, SDK, framework, version, or platform behavior, apply the web-search evidence gate before the claim.

Categories:

- Category A, specification definition: one official source may be enough when the claim is only what the spec says should happen.
- Category B, runtime behavior: run at least three WebSearch queries.
- Category C, version-dependent behavior: run at least three WebSearch queries, including the version or year.

Minimum query pattern for Category B or C:

- `<tool> <year> github issue`
- `<tool> reddit`
- `<tool> not working <version>`

Evidence block extension:

```text
- web-search-evidence:
  - query: <query 1>
    accessible_url: <url>
    finding: <key finding or value>
    source-locator:
      source_type: web
      region: n/a
  - query: <query 2>
    accessible_url: <url>
    finding: <key finding or value>
    source-locator:
      source_type: web
      region: n/a
  - query: <query 3>
    accessible_url: <url>
    finding: <key finding or value>
    source-locator:
      source_type: web
      region: n/a
```

Source-locator contract:

- Web evidence must include `accessible_url`.
- Attached or local file evidence must include `file_path`, `page`, and `region`.
- `region` values are `top`, `middle`, `bottom`, or `n/a`. Literal enum form: `top | middle | bottom | n/a`.
- Materials without pages use `page: n/a` plus an equivalent locator such as section, row, slide, or sheet in `locator_note`.
- Numeric claims, original sources, tables, and figures must bind the specific value to its source location.

When Category B or C appears and `web-search-evidence` has fewer than three entries, lacks `accessible_url`, or lacks `source-locator`, the completion claim is invalid. Search again, fill the evidence, then claim only what the evidence supports.

This gate exists because official docs describe intended behavior, while community reports often reveal runtime regressions, race conditions, and version-dependent failures.

The only exception is an explicit user instruction for this session to waive web-search evidence.

## Evaluator Artifact Contract

Before claiming verification-complexity-level-3 completion, external agent governance absorption, or RAG/evaluator candidate promotion, read `docs/policies/evaluator-artifact-contract.md`.

The completion evidence must include an accepted `verifier-result.json`.

- A read-only evaluator pass must not modify installed assets.
- Do not promote a candidate playbook without an accepted verifier result.
- At least one rejected candidate must exist so the verifier has proven it can say no.

## When To Apply

Apply this skill immediately before:

- any completion or success claim
- any recommendation or choice that claims finished work or verified results
- any new current-turn positive status judgment
- commit, push, PR creation, or branch finishing
- endorsing a delegated agent result as current after relevant state may have changed
- reporting tests, lint, build, scans, or review as sufficient

The rule covers exact words, paraphrases, implications, and tone that suggests the work is complete.

## Final Self-Check

Before finalizing, ask:

- What are the acceptance criteria?
- Which closure claims am I about to make?
- Does each new closure claim require fresh evidence, or is relevant unchanged evidence sufficient?
- What live uncertainty and next decision can the check change?
- Did I read the full output and exit status?
- Is anything still unverified?
- Does any claim require web-search evidence or an evaluator artifact?

Map the claim, apply the uncertainty gate, run a check only when its outcomes can change the decision, read the output, then speak.

Attribution

AidALLAidALL
View sourceMore from AidALL →
SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments (0)

No comments yet. Be the first to comment!

SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Related Skills

Caveman

Ultra-compressed communication mode that cuts output tokens while keeping technical accuracy. Levels: lite, full, ultra and the wenyan variants. Use for /caveman, "caveman mode", "talk like caveman", "be brief" or "less tokens".

1074701 votes

Hyperplan

Adversarial multi-agent planning skill. Self-orchestrates 5 hostile category members (unspecified-low, unspecified-high, deep, ultrabrain, artistry) via team-mode for ruthless cross-critique debate, distills only the defensible insights, then MANDATORILY hands the distilled insight bundle to the `plan` agent for executable plan formalization. Use when planning needs maximum rigor and surfacing of weak assumptions, blind spots, and over-engineering. Triggers: 'hyperplan', 'hpp', '/hyperplan', ...

695601 votes

Mcp Code Execution

Routes multi-tool workflows through MCP servers for large datasets and pipelines. Use when Bash tool overhead is limiting throughput on data-heavy tasks.

3351 votes

catchup

Recovers the conversation and failed tool calls of a previous Codex, Claude Code, Antigravity, Cline, Copilot CLI, Cursor, DeepSeek Harness, Kimi, OpenCode, Pi Agent, or ZCode session. Use when the user says "catch up", "what did the last session do", "get me up to speed", "I switched agents", asks to recover/summarize a previous session before continuing, or asks to diagnose or report a catchup failure. Do NOT use for the current conversation, git history, or any non-agent log.

691 votes

math-skill

A comprehensive mathematical reasoning skill for AI assistants — handles arithmetic to research-level problems with rigorous step-by-step reasoning, systematic verification, and transparent uncertainty handling

381 votes
View all in ai-agents →