Use before claiming completion, fixes, passes, commits, or PR creation. Requires running verification commands and reading their output before making success claims. Evidence always comes before claims.
Scanned 9/28/2026
Install to Claude Code
npx -y skills add AidALL/ghost-alice --skill verification-before-completion --agent claude-codeInstalls into .claude/skills of the current project.
Are you the author of Verification Before Completion?
Add the live security badge to your README — it updates automatically with every re-scan.
[](https://www.skillsdirectory.com/skills/aidall-verification-before-completion)More formats (shields.io, HTML) on the badges page.
---
name: verification-before-completion
description: Use before claiming completion, fixes, passes, commits, or PR creation. Requires running verification commands and reading their output before making success claims. Evidence always comes before claims.
compatibility:
- "Python 3.11+ standard library"
---
# Verification Before Completion
## Contents
- [Overview](#overview)
- [Iron Law](#iron-law)
- [Acceptance Criteria Iron Law](#acceptance-criteria-iron-law)
- [Relayed Verdicts And Absence Claims](#relayed-verdicts-and-absence-claims)
- [Hard Finalization Order](#hard-finalization-order)
- [Autopilot Proof Publication](#autopilot-proof-publication)
- [Retain Evidence On First Execution](#retain-evidence-on-first-execution)
- [Gate Function](#gate-function)
- [Evidence Selection And Stop Gate](#evidence-selection-and-stop-gate)
- [Completion-Check Format](#completion-check-format)
- [Common Failures](#common-failures)
- [Red Flags](#red-flags)
- [Rationalization Defense](#rationalization-defense)
- [Verification Patterns](#verification-patterns)
- [Why It Matters](#why-it-matters)
- [External Tool Web-Search-First Gate](#external-tool-web-search-first-gate)
- [Evaluator Artifact Contract](#evaluator-artifact-contract)
- [When To Apply](#when-to-apply)
- [Final Self-Check](#final-self-check)
## Overview
A new closure claim without decision-relevant fresh evidence is not efficiency. It is a lie.
Core principles:
- Evidence always comes before the claim.
- The letter of the rule and the spirit of the rule are the same rule.
- A surrounding signal is not direct proof unless it satisfies the relevant criterion.
## Iron Law
```text
Do not claim a new current-turn closure without decision-relevant fresh verification from this turn.
```
Explaining unchanged prior work is not a new closure claim. Cite the existing evidence and its age instead of rerunning unchanged work merely because another message arrived.
Every user input reopens routing; it does not by itself invalidate unchanged evidence or require reverification. Reverify when the relevant state, artifact, or criterion changed; a new error, mismatch, contradiction, or instability appeared; or the user explicitly requested a new check.
If a verification command or inspection did not run in this message, do not claim that it freshly passed in this message.
## Acceptance Criteria Iron Law
```text
No acceptance-criteria means no completed verification-before-completion.
```
Before any claim that executed work is complete, fixed, successful, or freshly verified, extract verifiable criteria from the user intent, locked decisions, and boundary-contract. Put those criteria in `acceptance-criteria`, then connect each intended closure claim to a criterion and fresh evidence in `claim-evidence-map`.
Evidence such as link checks, lint, diff checks, or passing tests proves completion only when it directly satisfies the criterion. If the central criterion is not directly verified, leave it in `unverified` and report partial status in prose.
## Relayed Verdicts And Absence Claims
A verdict you endorse as current is your own claim. Put it in the `claim-evidence-map`. If the relevant state may have changed, gather fresh evidence; inheriting a source's verdict is not evidence. If the task is only to explain an unchanged prior result, cite the existing evidence and its age without recreating it. Severity does not lower the bar.
An absence claim -- "no test exists", "X is not enforced", "nothing handles this" -- is never proven by reasoning or by a source's say-so. For a new or possibly changed absence claim, use a targeted current search that would surface the thing if present. Reuse a relevant prior search only when the searched state and criterion are unchanged, and state its age.
A verdict stated only in prose, outside the `claim-evidence-map`, escapes this gate. If you assert it, map it.
## Hard Finalization Order
Hard sequence for a new current-turn closure claim: skill load/call -> decision-relevant fresh verification -> [completion-check]
Before any executed-work completion, fix, success, or fresh-verification claim, perform the steps below in this exact order:
1. Load or call `verification-before-completion` for the current turn. On Claude Code, this means the visible Skill call. On Codex, this means reading this current `SKILL.md` and following its workflow.
2. Extract the acceptance criteria and run the decision-relevant fresh verification that can prove or disprove each intended final claim.
3. Only after the skill is loaded and the fresh evidence is read, write `[completion-check]` with `skill-call: verification-before-completion (this turn)`.
If any step is missing or out of order, the completion-check is invalid.
## Autopilot Proof Publication
When the current installation is Codex or Claude, includes `autopilot-mode/scripts/autopilot_completion.py`, and this session has admitted criteria for authorized execution, the installed PreToolUse adapter captures immutable prospective provenance before the business verification and surfaces one concrete preparation command for that current contract without admitting or advancing execution. Other platforms retain the existing workflow. Plan-only replies, explanations, installations without this helper, and sessions without admitted execution criteria do not use this publication path. Do not create execution criteria to activate it.
1. Keep the current admitted criteria and semantic contract accurate before verification. The pretool notice supplies current coordinates; follow its preparation command before the first final answer, preferably before verification. Repeated tools under the same contract do not repeat the notice or refresh its original capture time. Resolve the installed helper and the current hook/intake coordinates. Use the concrete notice command, or run `autopilot_completion.py prepare --reapprove-current-input --intent-root ROOT --platform PLATFORM --session-id SESSION --input-event-id INPUT` with the permitted Python interpreter. This explicit preparation keeps the admitted user task as the publication unit even when advisory conduct feedback exists; it does not replace that task with a separate conduct plan. Preserve the returned `receipt_token` and `prepared_at` strings verbatim. Use its returned `criterion_ids` unchanged for the current proof; historical met criteria keep their original evidence. Preparation promotes an exact existing prospective capture when available; otherwise it captures runtime provenance at that moment. It does not verify business work. For a newer input or changed contract that the user already authorized, use `prepare --reapprove-current-input` with that current input receipt. It can run after verification only when the runtime already captured the exact current contract before that original proof; the original capture and proof times remain unchanged. Without such a capture, prepare before its new verification. This uses the supported admission bridge and archives the previous generation while admitting the new current generation; it does not require another permission round for already-authorized work.
2. Perform the necessary business verification once. Capture its actual ISO timestamp as a string and its tool-result or evidence locator. Keep the original reference in the supported `[completion-check]` covering exactly the returned `criterion_ids`, with exactly one nonempty top-level `- evidence:` section. Supply the timestamp once through `--verified-at`; do not insert it into proof that lacks a timestamp. Preserve the original proof bytes and raw timestamp; never replace the verification time with the later publication time. If proof already declares `verified-at` or `verified_at`, its single value must match that raw string exactly. Existing explicit verification times in legacy evidence remain checked; an inconsistent, malformed, or ambiguous declaration is rejected.
3. Before the initial final answer, use the concrete publish command returned by preparation, or run `autopilot_completion.py publish --receipt-token TOKEN --verified-at ORIGINAL_ISO_TIME --completion-file -`, passing that exact completion block on stdin. The helper derives the exact reference text from the single top-level evidence section and binds the original raw timestamp, receipt and unchanged proof digest in immutable publication provenance. This checks consistency, not the truth of a caller's first supplied timestamp. It does not select nested claim evidence, invent a tool identifier, or independently authenticate an external tool result. Missing, empty, duplicate, placeholder, ambiguous, or partial evidence remains rejected. A legacy caller may explicitly supply `--evidence-source` only with an unchanged value already present in the proof; a mismatch or empty override is rejected, never replaced. Both helper invocations must use the current session's `GHOST_ALICE_SESSION_INTENT_ROOT`, `GHOST_ALICE_PLATFORM`, and `GHOST_ALICE_SESSION_ID` environment bindings; derive these from the current hook/intake, never from an older run. The helper constructs the decision envelope and digest. Its pending-publication result is bookkeeping evidence; the Stop adapter retains the transaction that marks criteria met. Keep the requested business answer and supported completion block in the final response.
Do not hand-build a replacement decision envelope, rerun business checks to manufacture a publication, or prepare a new receipt for proof produced before its capture. If publication alone failed, retry publication with the original receipt and unchanged proof. A missing publication record is not evidence of unfinished business work. A stale receipt, changed input, changed criterion, changed scope, or failed proof stays rejected; reopen only the affected verification under the current authorized contract. At Stop, a matched prospective capture becomes a receipt and the missing-publication message supplies the concrete installed publish command. Follow that command with the original proof and timestamp. If no original receipt or matching prospective capture exists, report the binding gap separately from the supported business result instead of relabeling old proof. Use stdin or permitted runtime scratch when the user prohibits extra task files.
## Retain Evidence On First Execution
Before running a necessary check or inspection, decide how to retain its returned output, exit status, source locator and actual verification time. In an orchestrator call, keep the original result in session state before returning it when later steps need that object. This avoids losing a successful result at an execution boundary.
An already returned successful tool result is evidence. Cite its existing locator or copy its unchanged contents into permitted scratch if persistence is needed. Do not rerun the same successful inspection merely to populate `store`, add redirection, create an evidence file, recover an execution-local variable, or satisfy publication bookkeeping. Missing storage is not missing verification. Rerun only when the original result is unavailable or incomplete, the relevant state changed, or a decision-relevant contradiction requires it; state that reason before running.
Preserve the original source and time when copying evidence. Do not manufacture a fresh verification timestamp while saving it. If the source or original time cannot be established, report that specific evidence gap rather than relabeling an old result. Publication remains required before the first final answer for the admitted execution path; retaining evidence does not replace publication or establish a verdict.
## Gate Function
Before claiming any state as satisfied:
1. Criterion: extract `acceptance-criteria` from the user intent and contract.
2. Mapping: connect each claim you plan to make to one criterion.
3. Uncertainty: name the live uncertainty and the next decision each possible outcome can change.
4. Evidence target: identify the command, file, source locator, or tool output that can prove each criterion.
5. Execution: run the smallest decision-relevant check when the uncertainty gate requires fresh evidence.
6. Reading: read the full output, exit code, and failure count.
7. Judgment: decide whether the output supports the criterion and claim.
8. Unverified handling: keep any unsupported criterion in `unverified`.
9. Claim: state only the range that was actually verified.
Skipping any step is not verification. It is a lie.
## Evidence Selection And Stop Gate
Current accessible behavior or content is the default direct evidence for semantic claims. Hash or provenance evidence is appropriate when the criterion is artifact identity, integrity, drift, merge safety, or reproducibility. Do not use hash equality, byte identity, cache history, or repository lineage as a proxy for current semantic behavior.
Before running a check, name the live uncertainty and the next decision that each possible outcome can change. If no possible outcome can change the criterion or next decision, do not run the check.
Verification output does not create a new obligation to verify the verification. Stop when a repeated check produces no relevant state delta, or when further checking would displace the user's primary objective. Resume only after a state change, new error, mismatch, contradiction, instability, or explicit request.
## Completion-Check Format
Use this block immediately before the final summary when you are making an executed-work completion, fix, success, or fresh-verification claim.
```text
[completion-check]
- verification-before-completion: done
- skill-call: verification-before-completion (this turn)
- acceptance-criteria:
- <criterion-id>: <user-intent-or-contract-condition> [source: user-explicit | inferred | previous-tool | system-doc]
- claim-evidence-map:
- claim: <completion-or-recommendation-claim>
criterion: <criterion-id>
evidence: <fresh command, inspected file, source locator, or tool output>
verdict: pass | fail
- unverified:
- none
- evidence: <fresh command or inspected file>
```
Serialize `claim`, `criterion`, `evidence`, and `verdict` on their own physical lines. Emit an evidence-supported bare `pass` or `fail` verdict with no trailing punctuation, quotes, markup, or explanatory prose. If evidence does not support a verdict, report honest partial state without a finalized `[completion-check]`. Record the actually called `verification-before-completion` skill in an explicit `skills-loaded` list in `[io-trace]`. A format repair must preserve the substantive business result and supported evidence.
Only emit a finalized `[completion-check]` when every listed criterion has a `pass` or `fail` verdict and `unverified` is `none`. If anything remains unverified, do not emit the final block. Report the partial state in prose and name the missing check.
## Common Failures
| Claim | Required evidence | Insufficient evidence |
| --- | --- | --- |
| Tests passed | Fresh test command output with zero failures | A previous run or a prediction |
| Lint is clean | Fresh lint output with zero errors | Partial lint or an extrapolation |
| Build succeeded | Build command exit code 0 | Lint passing |
| Bug fixed | A test or reproduction that covers the original symptom | Changed code plus confidence |
| Regression test works | Red-green evidence when TDD requires it | A test that passed once |
| Agent completed the work | VCS diff plus independent verification | The agent's success report |
| Requirements satisfied | Claim-evidence map for each acceptance criterion | Tests pass alone, links pass alone, or diff exists alone |
| Relayed/endorsed review verdict | Current behavior evidence when state may have changed; otherwise the relevant existing evidence with its age | The reviewer's verdict, or your agreement with it, alone |
| Absence claim ("no test/code exists", "not enforced") | A targeted current search when absence may have changed; otherwise the relevant existing search with its age | Reasoning or the source's say-so |
## Red Flags
Stop before claiming success when any of these appear:
- "should", "probably", or "seems to"
- satisfaction language before verification
- a new closure, commit, push, or PR claim without decision-relevant checks
- trusting another agent's success report
- relying on partial verification
- wanting to finish because the work feels close
- treating lint, diff, or tests as completion without criterion mapping
- implying success through wording while avoiding the word "done"
## Rationalization Defense
| Excuse | Required response |
| --- | --- |
| "It should work now." | Run verification. |
| "I am confident." | Confidence is not evidence. |
| "Just this once." | No exception. |
| "Lint passed." | Lint is not a compiler or a requirement map. |
| "Another agent said it succeeded." | Endorse it only with decision-relevant evidence; reuse unchanged evidence with its age. |
| "Partial checks are enough." | Partial checks prove only the checked criteria. |
| "The wording is different, so the rule does not apply." | Completion implications still count. |
## Verification Patterns
Tests:
```text
Run the test command, read the result, then claim only the observed result.
```
Regression tests:
```text
Write the test -> run and observe pass -> revert or disable the fix -> observe fail -> restore fix -> observe pass.
```
Builds:
```text
Run the build command and read exit code 0 before claiming build success.
```
Requirements:
```text
Re-read the user intent and contract -> write acceptance criteria -> verify each criterion -> report missing criteria or verified completion.
```
Agent delegation:
```text
Read the agent report -> inspect current accessible behavior when needed -> apply the uncertainty gate -> run only a decision-relevant check -> report the supported state.
```
## Why It Matters
From accumulated failure memory:
- A user said "I cannot trust you" and trust broke.
- An undefined function shipped and a crash followed.
- A missing requirement shipped as an incomplete feature.
- False completion wasted time, forced a change of direction, and caused rework.
- The standing rule for a violation is this. Honesty is a core value. If you lie, you are replaced.
## External Tool Web-Search-First Gate
Layer marker: `web-search-first`.
If the final claim includes factual behavior about an external tool, library, CLI, SDK, framework, version, or platform behavior, apply the web-search evidence gate before the claim.
Categories:
- Category A, specification definition: one official source may be enough when the claim is only what the spec says should happen.
- Category B, runtime behavior: run at least three WebSearch queries.
- Category C, version-dependent behavior: run at least three WebSearch queries, including the version or year.
Minimum query pattern for Category B or C:
- `<tool> <year> github issue`
- `<tool> reddit`
- `<tool> not working <version>`
Evidence block extension:
```text
- web-search-evidence:
- query: <query 1>
accessible_url: <url>
finding: <key finding or value>
source-locator:
source_type: web
region: n/a
- query: <query 2>
accessible_url: <url>
finding: <key finding or value>
source-locator:
source_type: web
region: n/a
- query: <query 3>
accessible_url: <url>
finding: <key finding or value>
source-locator:
source_type: web
region: n/a
```
Source-locator contract:
- Web evidence must include `accessible_url`.
- Attached or local file evidence must include `file_path`, `page`, and `region`.
- `region` values are `top`, `middle`, `bottom`, or `n/a`. Literal enum form: `top | middle | bottom | n/a`.
- Materials without pages use `page: n/a` plus an equivalent locator such as section, row, slide, or sheet in `locator_note`.
- Numeric claims, original sources, tables, and figures must bind the specific value to its source location.
When Category B or C appears and `web-search-evidence` has fewer than three entries, lacks `accessible_url`, or lacks `source-locator`, the completion claim is invalid. Search again, fill the evidence, then claim only what the evidence supports.
This gate exists because official docs describe intended behavior, while community reports often reveal runtime regressions, race conditions, and version-dependent failures.
The only exception is an explicit user instruction for this session to waive web-search evidence.
## Evaluator Artifact Contract
Before claiming verification-complexity-level-3 completion, external agent governance absorption, or RAG/evaluator candidate promotion, read `docs/policies/evaluator-artifact-contract.md`.
The completion evidence must include an accepted `verifier-result.json`.
- A read-only evaluator pass must not modify installed assets.
- Do not promote a candidate playbook without an accepted verifier result.
- At least one rejected candidate must exist so the verifier has proven it can say no.
## When To Apply
Apply this skill immediately before:
- any completion or success claim
- any recommendation or choice that claims finished work or verified results
- any new current-turn positive status judgment
- commit, push, PR creation, or branch finishing
- endorsing a delegated agent result as current after relevant state may have changed
- reporting tests, lint, build, scans, or review as sufficient
The rule covers exact words, paraphrases, implications, and tone that suggests the work is complete.
## Final Self-Check
Before finalizing, ask:
- What are the acceptance criteria?
- Which closure claims am I about to make?
- Does each new closure claim require fresh evidence, or is relevant unchanged evidence sufficient?
- What live uncertainty and next decision can the check change?
- Did I read the full output and exit status?
- Is anything still unverified?
- Does any claim require web-search evidence or an evaluator artifact?
Map the claim, apply the uncertainty gate, run a check only when its outcomes can change the decision, read the output, then speak.
Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.
No comments yet. Be the first to comment!