Skills DirectorySkills Directory
SkillsLearnSecurityCategoriesDocsCommunityBlog
Sign InSubmit Skill
Skills Directory

Security-tested agent skills for Claude, coding agents, and AI workflows.

Directory

  • Browse Skills
  • All Skills A–Z
  • Claude Skills
  • Claude Code Skills
  • Agent Skills
  • Categories
  • Authors
  • Submit a Skill

Learn

  • Learn Hub
  • Install Claude Skills
  • Write SKILL.md
  • Skills vs MCP
  • Directories Compared

Security

  • Security
  • Methodology
  • Secure Claude Skills
  • Security Badges

Company

  • About
  • Community
  • Blog
  • API Docs
  • Advertise

2026 Skills Directory. All rights reserved.

ProTermsPrivacyRefunds
Back to skills

Ship Check

ASecurity

The definition of done. Run before calling any change done, fixed, verified, ready or shippable, for features and bug fixes alike. Checks the pre-mortem was answered, proves the tests fail on the old code, gets a fresh breaker review, runs CI's own checks in a clean checkout, checks the user-facing words, and writes the evidence report. Also use when asked "is it done", "is it ready", "did you verify it" or "can we merge".

44 stars
0 votes
0 copies
0 views
Added 9/28/2026
ai-agentsgogitdatabasesecurity

Works with

claude codecursor

Security Analysis

A100/100

Scanned 9/28/2026

Install to Claude Code

$npx -y skills add joetawil7/first-pass --skill ship-check --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Ship Check?

Add the live security badge to your README — it updates automatically with every re-scan.

Security grade badge for Ship Check
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/joetawil7-ship-check/badge)](https://www.skillsdirectory.com/skills/joetawil7-ship-check)

More formats (shields.io, HTML) on the badges page.

Files
SKILL.md
---
name: ship-check
description: The definition of done. Run before calling any change done, fixed, verified, ready or shippable, for features and bug fixes alike. Checks the pre-mortem was answered, proves the tests fail on the old code, gets a fresh breaker review, runs CI's own checks in a clean checkout, checks the user-facing words, and writes the evidence report. Also use when asked "is it done", "is it ready", "did you verify it" or "can we merge".
---

# ship-check

Walk every step. A step you cannot do goes in the report under "Not verified" with the
reason; it is never skipped silently. The repo's commands, test limits and heavy-run rules
are in the `first-pass:project` block of its instruction file (AGENTS.md or CLAUDE.md); if
there is none, read them from the CI config and say so in the report. Its test limits bind
every step below.

Run everything inside the repo that changed (`cd <repo>`, `git -C <repo>`), not from a main
folder above it. A change that spans repos walks the steps once per repo.

A prompt with several items walks steps 1 and 2 per item, while each is built, running only
the tests that item touches; one clean checkout of the base serves every item's fail-first
run. Steps 3 and 4 run once for all of them, after the last item is built: one review per
item (small items that touch the same code can share one), started together only where the
repo's test limits say side-by-side runs are safe (at most three at once), otherwise one
after another; then CI's full checks once, in one clean checkout holding every item. The
report answers each item.

Nothing is pushed, merged, deployed, migrated or published until steps 3 and 4 are finished
for a clean checkout holding exactly what it ships (failures the base has too are named,
and findings left open are answered by the user first), unless the user says to ship it as
it is.

## 0. Size

A change with no logic in it (a comment, a doc, a spelling fix that changes no behaviour)
runs only the repo's format, lint and build checks, plus step 6 when a person reads the
text, and the report says this exception was used. Anything else, however small (a
constant, a condition, a default, a price, a label whose meaning changes), walks every step.
When unsure, it is not a typo.

## 1. Pre-mortem

Find the pre-mortem in the plan. If there is none, write it now from the diff with the
`premortem` skill, and say in the report that it was written after the code. Search again
for neighbors of every field, status, option, queue and endpoint in the diff; add any the
plan missed.

## 2. Tests that fail on the old code

For each behaviour the change adds or fixes, and each pre-mortem answer of the "test" kind:

1. **Right layer.** The lowest layer that reproduces it end to end. If the bug could live
   in a query, a transaction, a queue or a browser, the test runs against the real thing
   (an integration test on a real database, an end-to-end test in a browser). A unit test
   with mocks is enough only for pure logic.
2. **Fails first.** In a clean checkout of the code before the change. That is the base:
   HEAD when the change is still uncommitted, otherwise the commit the branch started from
   (`git merge-base HEAD origin/<main branch>`). Never a checkout that already contains
   the change, or good tests pass on the "old" code and look worthless.
   ```
   git -C <repo> worktree add --detach "<temp dir>/<repo>-shipcheck-<short base sha>-<time>" <base>
   ```
   Put it in the system temp folder, not beside the repo (a new folder inside a main folder
   looks like a new repo to every tool that scans it), under a name no parallel session
   will pick.
   Copy in only the new or changed test files, install dependencies as CI does, run them,
   and record which fail and why. Then copy in the change and record that they pass. A test
   that passes on the old code proves nothing about the change: rewrite it.
3. **Guards** (tests that behaviour which must not change still holds) pass on both. Say
   which tests are guards.
4. **The risky answers get a test.** When the change has a queue job, a paid call, a
   publish, a payment or a delete in its path: one test runs it twice at once, one makes the
   outside call time out or return a 5xx.
5. **Tests assert what the user sees or what is stored**, not only that a new test id
   exists or that a mock was called.

## 3. Fresh review

Hand the change to the `breaker` agent (Claude Code: `first-pass:breaker` from the plugin,
or `breaker` where a repo installed its own; Cursor: `/breaker`) with what the change is
for, the repo, the base ref or file list, and the pre-mortem. It must run in its own
context. If your tool cannot start one, ask the user to run the breaker in a new chat;
never review in the context that wrote the code.

For each finding:

- Real (CONFIRMED, or PLAUSIBLE and you confirm it), and its scenario breaks the task or a
  promise in the repo's rules, invariants or docs, or does real harm (money lost, wrongly
  charged or spent without a cap, lost or leaked data, a side effect done twice, a security
  hole, a legal breach, a crash): fix it,
  with its own failing-first test (step 2).
- Real, but its worst case stays inside what the repo promises: list it under Open with
  why; don't build for it.
- Disagree: say why in the report, with file:line.
- Real but out of scope (the change neither caused it nor made it worse): list it under
  Open.

If the fixes were more than small, run the breaker again on the fixes: round 2, and round 3
on round 2's fixes if they were more than small too. A fix that touches code another item
in the same prompt uses always gets round 2, on the combined diff. Rounds are counted per
item; say the count in the reply after each round ("review round 2 of 3 for item 1"), so
it survives a compacted context. After round 3, stop: only a finding that does real harm
(see above), whatever severity it was given, is still fixed, and so is a CI failure the
change caused (step 4); each such fix gets a review of that fix, repeated until one finds no
new real harm in it. Disputed and out-of-scope findings stay listed, not fixed again. Every
other finding goes under Open with its worst case, for the user to decide, and an item left
with an open finding in its scope that breaks the task or a promise in the repo's rules,
invariants or docs is reported as built, not done. More rounds for other findings only when
the user asks.

## 4. CI's own checks, in the clean checkout

In the step 2 worktree with the whole change copied in, run exactly the commands CI runs
(format, lint, typecheck, build, unit tests, integration and end-to-end tests the change
touches; the whole integration suite when the change touches jobs, payments, publishing,
deletion or auth). Follow the project's rules for heavy runs.

A failure that also happens on the base without the change is pre-existing: name the test
and move on. A failure the change caused gets a fix; a fix that is more than small goes
back to step 3 (it counts as a round; after round 3 it is reviewed like a fix for real
harm), and the full checks run again. Then remove the worktree (`git -C <repo> worktree remove --force <path>`) and any leftovers,
including copied env files.

## 5. Monitoring

List what now reaches monitoring and at what level. Every error the change swallows is
reported at a level someone will see. For a bug fix, consider a tripwire: a report that
fires if the exact failure ever happens again.

## 6. Words

Search UI strings, emails and notifications, help, docs, and pricing and legal pages for
sentences about what changed. Each is confirmed true or changed. Changes to legal, pricing
or public copy get their own line in the report.

## 7. Invariants

For each invariant in `INVARIANTS.md` the change touches: held, or broken (add it to Known
breaks). If the change fixes a known break, move its id out.

## 8. Report

In plain words. Leave out any line with nothing in it, except Public copy changed and Fresh
review.

```
<What changed, one plain line: what the user can now do or will notice>
Verified: <what was run and how much of it, said plainly> → <result>, one line each (the test that failed before and passes now, with pass and fail counts; CI's checks)
Fresh review: <what the second reviewer found: n fixed, n disputed, n open>
Not verified: <each thing, and why>
Not handled, because: <each, from the pre-mortem>
Not built: <each guess left out, one line each>
Public copy changed: <file, or "none">
Open: <follow-ups, one line each>
```

"Verified" only ever sits next to something run in this session and its result. The exact
command and file:line go in when the user will use them, when something failed, for each
disputed finding (step 3), and wherever a rule or skill asks for them. If any step above was
skipped, the change is not done: say "not done" and why, not "done with caveats".

Attribution

joetawil7joetawil7
View sourceMore from joetawil7 →
SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments (0)

No comments yet. Be the first to comment!

SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Related Skills

Caveman

Ultra-compressed communication mode that cuts output tokens while keeping technical accuracy. Levels: lite, full, ultra and the wenyan variants. Use for /caveman, "caveman mode", "talk like caveman", "be brief" or "less tokens".

1074701 votes

Hyperplan

Adversarial multi-agent planning skill. Self-orchestrates 5 hostile category members (unspecified-low, unspecified-high, deep, ultrabrain, artistry) via team-mode for ruthless cross-critique debate, distills only the defensible insights, then MANDATORILY hands the distilled insight bundle to the `plan` agent for executable plan formalization. Use when planning needs maximum rigor and surfacing of weak assumptions, blind spots, and over-engineering. Triggers: 'hyperplan', 'hpp', '/hyperplan', ...

695601 votes

Mcp Code Execution

Routes multi-tool workflows through MCP servers for large datasets and pipelines. Use when Bash tool overhead is limiting throughput on data-heavy tasks.

3351 votes

catchup

Recovers the conversation and failed tool calls of a previous Codex, Claude Code, Antigravity, Cline, Copilot CLI, Cursor, DeepSeek Harness, Kimi, OpenCode, Pi Agent, or ZCode session. Use when the user says "catch up", "what did the last session do", "get me up to speed", "I switched agents", asks to recover/summarize a previous session before continuing, or asks to diagnose or report a catchup failure. Do NOT use for the current conversation, git history, or any non-agent log.

691 votes

math-skill

A comprehensive mathematical reasoning skill for AI assistants — handles arithmetic to research-level problems with rigorous step-by-step reasoning, systematic verification, and transparent uncertainty handling

381 votes
View all in ai-agents →