Skip to content
Back to skills

Cage Review

ASecurity

Review the contracts of a cage project — judge whether the tests really check what each invariant promises — and record the verdict so that `cage check` is clean. Use when `cage check` or the Stop hook reports REVIEW_MISSING, REVIEW_STALE or REVIEW_WEAK, or when asked to review a design.

  • 2 stars
  • 0 votes
  • 0 copies
  • 0 views
  • Added October 5, 2026
developmentgodatabase

Security analysis

A100/100

Scanned October 6, 2026

npx -y skills add vitalik1921/cage --skill cage-review --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Cage Review?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Cage Review
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/vitalik1921-cage-review/badge)](https://www.skillsdirectory.com/skills/vitalik1921-cage-review)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: cage-review
description: Review the contracts of a cage project — judge whether the tests really check what each invariant promises — and record the verdict so that `cage check` is clean. Use when `cage check` or the Stop hook reports REVIEW_MISSING, REVIEW_STALE or REVIEW_WEAK, or when asked to review a design.
---

# cage-review

`cage check` sees that a test is *tagged* with an invariant. It cannot see whether the test *checks* it. That is this review: read the design, the code and the tests of a contract together and say, per invariant, whether the tests would catch a broken promise. The verdict is recorded against a fingerprint of the material; `cage check` then requires a fresh one.

## Steps

1. `cage review` prints an index: every contract that needs a review (no verdict yet, or the material changed since), what changed and which invariants it touches. Take one at a time: `cage review Name` prints its material — for an outdated review the lines that changed and the previous finding of each invariant, the rest by file and line (`--files all` for every file whole, `--files none` for references only); `--format json` for the same with a JSON Schema of the verdict.
2. The packet is facts, without instruction: this skill is the instruction. Read the whole packet before judging: the design (business rules and invariants), the implementations, the test files, the helpers loaded with them, the `check:` codes (`cage codes` explains one), the files **used outside the module** (they rely on the contract's promises: a change to an invariant reaches them, and they may assume the old one — say so in a contract-level finding), and the list of files that are **not loaded** — open those in the repository; if you cannot, the assessment for what depends on them is `insufficient-context`, not a guess.
3. For every invariant, with the tests tagged `@covers` for it in front of you, answer the same questions the packet's instruction lists. When the review is outdated, judge afresh the invariants the packet marks as touched; for the others confirm the previous finding, or revise it if it no longer holds — every invariant still gets a finding:
   - Does the test exercise the behaviour the invariant is about, or the right external scenario?
   - Would it fail if exactly this promise were broken, and only then?
   - Do the assertions observe the right result, order or effect — not just that nothing was thrown?
   - Do mocks or stubs stand in for the very guarantee the test claims to check? The rule: a test that asserts the exact statement, clause or call that carries the promise is `adequate` for that clause; a clause left unasserted while a stub returns a canned result regardless is `weak`; what only a real database or service can show (one row kept, an id unchanged, a value the column type rejects) is `weak` even when the clause is asserted, unless a test against the real thing is linked — name the test to tag. The packet lists the module's test files that carry no tag: proof is often there, one tag away. The helpers the tests import are loaded with them: read the stub before judging what a test observes.
   - A boundary the contract's types allow counts even when today's callers cannot reach it (`""` where the type says `string`, `undefined` where it says `Slot`): the contract is the promise, not the callers. Say so in the reason; the design's owner may narrow the type instead.
   - Are the relevant errors, boundaries, concurrency and retries covered?
   - For a rule about the whole interface: is the interaction of methods checked?
   - Where there are several implementations, does each have the scenarios it needs?
   - Does the implementation do more than the invariant says (an extra condition the tests never touch)?
   - Does the business description add material requirements, or contradict the invariants?
4. Write the verdict as JSON in the shape the packet ends with (`## Verdict`, the fingerprints filled in): one entry per contract with its `fingerprint` copied from the packet, and one finding per invariant — `assessment` is `adequate`, `weak`, `unrelated` or `insufficient-context`; `reason` says why, first sentence first (it is printed by `cage check`); `evidence` is `file:line` or `file:line-line` of the test or code the judgement rests on (several joined with `; `; null only for `insufficient-context`); `suggestedChange` is what would make it adequate — or, for an adequate finding, an improvement worth making, or null. The packet's files are printed with line numbers. A contract without invariants gets one finding with `invariant: null`. Contract-level findings (`invariant: null`, as many as you have) may also stand next to the per-invariant ones: that is where an observation about the design or the code goes that is not a test weakness (a constraint no rule mentions, a dead method, a rule the callers cannot reach, a file outside the module that assumes the old promise) — assessed `adequate` so that the check does not fail on it, with the observation as the reason, first sentence first; `--record` prints them and the next packet of the contract repeats them. Evidence may cite a file the packet does not hold when you opened it in the repository. A test that proves an invariant of this contract but is tagged for another one (an e2e test through a port) is not this contract's evidence until it is tagged for it: say which test to tag in `suggestedChange`. `Lock` in a contract's header is its change policy (`@final` / `@extendable`), not a review matter.
5. `cage review --record <file>`. It refuses a verdict for other material (the fingerprint changed: review again), an unknown invariant, an invariant left unassessed, or a finding — contract-level ones included — with a blank reason or without evidence (null is allowed only for `insufficient-context`).
6. `cage check`: every finding that is not `adequate` is reported at the invariant. A `weak` or `unrelated` finding is work: fix the test (or the design, if the invariant is unobservable), then review again — the fix changes the material, so the next `cage review` includes it.

## Rules

- You judge as a stranger would, also when you wrote the code or the tests under review. The verdict is recorded and read by others.
- `adequate` needs evidence: name the test and what it asserts. "The test exists" is not evidence.
- Never lower an assessment, remove or soften an invariant, or edit `.cage/review.json` by hand to make the check pass. `cage review --accept` is not a review: it records the material as accepted without a verdict, and only a person decides that; do not run it unless asked to.
- Do not run the project's tests to decide; this is a reading of what the tests would prove, not whether they pass today.
- The packet lists the fingerprinted dependencies; how far the fingerprint reaches is the project's `reviewDependencies` setting. Do not change `.cage/config.json` to widen or narrow it.

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…