Skills DirectorySkills Directory
SkillsLearnSecurityCategoriesDocsBlogPro
Sign InSubmit Skill
Skills Directory

Security-tested agent skills for Claude, coding agents, and AI workflows.

Directory

  • Browse Skills
  • All Skills A–Z
  • Claude Skills
  • Claude Code Skills
  • Agent Skills
  • Categories
  • Authors
  • Submit a Skill

Learn

  • Learn Hub
  • Install Claude Skills
  • Write SKILL.md
  • Skills vs MCP
  • Directories Compared

Security

  • Security
  • Methodology
  • Secure Claude Skills
  • Security Badges
  • Chrome Extension
  • Skill Manager

Company

  • About
  • Community
  • Blog
  • API Docs
  • Advertise

2026 Skills Directory. All rights reserved.

ProTermsPrivacyRefunds
Back to skills

Code Review Battery

ASecurity

Use when reviewing code changes — one combined reviewer by default, parallel specialized reviewers only on diff signals. Invoke as: /sp-cr-battery [min-score] [--security|--no-security] [--mode=bug-fix|feature] (optional 1.0–10.0 quality threshold, default 7.0; default 9.2 in Bug Fix Review Mode). Bug Fix Mode auto-activates on hotfix/* and fix/[A-Z]+-[0-9]+ branches.

8 stars
0 votes
0 copies
0 views
Added 10/6/2026
ai-agentsgoshellbashsqlcode-reviewgitapisecurityperformance

Works with

cliapimcp

Security Analysis

A100/100

Pro scans all 21 files and shows the line behind each finding

Scanned 10/6/2026

$npx -y skills add bordenet/superpowers-plus --skill code-review-battery --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Code Review Battery?

Add the live security badge to your README — it updates automatically with every re-scan.

Security grade badge for Code Review Battery
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/bordenet-code-review-battery/badge)](https://www.skillsdirectory.com/skills/bordenet-code-review-battery)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
Files
DESIGN.md
---
name: code-review-battery
description: "Use when reviewing code changes — one combined reviewer by default, parallel specialized reviewers only on diff signals. Invoke as: /sp-cr-battery [min-score] [--security|--no-security] [--mode=bug-fix|feature] (optional 1.0–10.0 quality threshold, default 7.0; default 9.2 in Bug Fix Review Mode). Bug Fix Mode auto-activates on hotfix/* and fix/[A-Z]+-[0-9]+ branches."
summary: Default is ONE combined reviewer (Defect Finder + Guardian + Standards Enforcer lenses, plus a mandatory placement question); up to 5 more specialists (Design Critic, Performance Analyst, AttackerPersona, ShellRuntimeAuditor, BugPath Verifier) are dispatched in parallel only on diff signals, with source context for ripple analysis. AttackerPersona is signal-driven (security-sensitive diffs) and toggleable via --security/--no-security. ShellRuntimeAuditor is signal-driven (shell content -- a shebang, a .sh/.bash file, or tool-wrapper code) with no manual toggle. Aggregates findings with triple-filter prioritization and Round 2 escalation.
triggers:
  - /sp-cr-battery
  - /sp-deepreview
  - code review battery
  - battery review
  - parallel review
  - dispatch reviewers
  - specialist review
  - multi-reviewer
  - review with battery
aliases: [CRB, review-battery]
anti_triggers:
  - requesting code review
  - receiving code review
coordination:
  group: code-quality
  order: 1
  enables:
    - progressive-code-review-gate
    - verification-before-completion
  requires: []
  escalates_to: []
  internal: false
source: superpowers-plus
augment_menu: true
composition:
  consumes: [code-changes]
  produces: [review-feedback]
  capabilities: [dispatches-review, parallel-review]
  priority: 25
---

# Code Review Battery

> **Mechanical routing:** don't decide from memory or from the "Wrong skill?" prose below -- run `tools/review.sh route <path> [<path> ...]` first (paths of the files you're about to review). It wraps `tools/which-gate.sh` and prints the correct skill + sentinel + runner for each artifact. If the router says a different skill, follow the router, not this banner. If the router errors or is unavailable, stop and report -- do not fall back to the prose. The banner is an inner backstop, not a substitute for the mechanical check.
>
> **Wrong skill?** File-protocol review handoff → `code-review`. PR inline → `providing-code-review`. Pre-commit gate → `progressive-code-review-gate`. Full-repo security audit → `repo-security-scan` or `/sp-devsec-audit`. Skill.md files or skill-adjacent tooling infrastructure → `llm-skill-review`. **Slash commands:** `/sp-cr-battery [min-score]` (primary; optional 1.0–10.0 quality threshold, default 7.0), `/sp-deepreview` (legacy).

Run ONE combined reviewer by default. Dispatches up to 7 specialized reviewer agents in parallel, but only those whose signal the diff carries. A triage coordinator selects which to activate based on the diff, then aggregates findings into a unified report.

**Why this exists**: A single reviewer tries to evaluate everything simultaneously, leading to shallow coverage, inconsistent focus, and ~40% false positive rates. Specialized reviewers with focused prompts produce deeper analysis with near-zero false positives.

## When to Use

- When `requesting-code-review` or `progressive-code-review-gate` triggers a review
- When you want a thorough review of staged changes, a commit range, or a PR diff
- When reviewing someone else's code
- When `verification-before-completion` detects the implementation→presentation transition and no valid sentinel exists for HEAD

**Gate chain position 3 of 4:**

| Gate (order) | Self-fires when | Short-circuit if |
|---|---|---|
| 1. `debate` | About to commit to a design before coding | Already ran this session |
| 2. `progressive-harsh-review` | About to present a non-code deliverable | Already ran on this artifact |
| **3. `code-review-battery`** | **About to present/commit/push code** | **Valid sentinel for HEAD exists** |
| 4. `verification-before-completion` | About to write any results-presenting response | Sentinel SHA == HEAD → skip re-dispatch |

**One-per-unit rule:** Battery fires at most once per coherent unit of work. If a valid `.code-review-cleared` sentinel exists for HEAD, the gate is already satisfied — do not re-dispatch.

## Procedure

### Phase 0: Sentinel Check (canonical skip gate — run before dispatching anything)

This is the canonical skip gate for the one-per-unit rule. **Run `tools/review-preflight.py` first** (`--base <integration-branch>`, default `origin/dev`, or `--staged` for an uncommitted change; pass `--mode=bug-fix|feature` when given) and use its JSON -- sentinel states, BugPath mode, fired signal rows with file:line hits, rows left to judgment, and the inline-exemption answer -- instead of applying the Phase 0-1 tables by hand; the tables stay the spec. Callers (`requesting-code-review`, `finishing-a-development-branch`, `progressive-code-review-gate`) should run this before dispatching. If a caller does not implement Phase 0 explicitly, the agent should run the preflight (or apply the tables) before invoking battery.

```bash
SENTINEL="$(git rev-parse --show-toplevel 2>/dev/null || echo .)/.code-review-cleared"
# cat succeeds on a stale sentinel (wrong SHA) — compare the printed SHA to HEAD explicitly.
cat "$SENTINEL" 2>/dev/null || echo "NO CLEARANCE"
echo "HEAD: $(git rev-parse HEAD 2>/dev/null)"
git diff --quiet && git diff --cached --quiet && echo "WORKTREE_CLEAN" || echo "WORKTREE_DIRTY"
```

| Sentinel state | Decision |
|----------------|----------|
| `NO CLEARANCE` | Run battery (proceed to Phase 1). |
| Sentinel SHA ≠ HEAD SHA, and preflight does not report it `carried` (no in-scope file changed) | Run battery (battery is stale). |
| Sentinel valid for HEAD but `WORKTREE_DIRTY` | Run battery (staged/unstaged changes exist that were not reviewed). |
| Valid sentinel for HEAD (or `carried`) AND `WORKTREE_CLEAN` | **Skip.** Battery already ran on the current code. Note the clearance and skip to Phase 6. |
| Malformed, or preflight `not-clearing` (verdict does not clear the gate) | Delete `.code-review-cleared`, run battery. |

### Phase 0.5: BugPath Mode Detection

> **Load gate:** This phase requires `reference.md` to be loaded alongside this skill. If absent AND the branch matches a BugPath pattern (`hotfix/*` or `fix/[A-Z]+-[0-9]+`, per the detection script in `reference.md`): **hard halt** — emit only: "Cannot enter BugPath Mode: reference.md is absent. Load it before proceeding." Output nothing further — no phase summaries, no partial results, no clarifying questions. Return control to the user. If absent AND the branch does NOT match a BugPath pattern: treat BugPath Mode as INACTIVE and note in triage line.

Run immediately after the sentinel check. Detect whether this is a targeted bug fix, then set the mode before triage.

(Detection script: `reference.md` § BugPath Detection Script.)

| Signal | BugPath Mode trigger |
|--------|---------------------|
| Branch prefix `hotfix/*` | Active |
| Branch matching `fix/[A-Z]+-[0-9]+` (e.g., `fix/PROJ-1234` or `fix/PROJ-1234-description`) | Active |
| Explicit flag `--mode=bug-fix` | Active |
| Explicit flag `--mode=feature` | Inactive (overrides branch detection) |

**When BugPath Mode is active:** BugPath Verifier mandatory (not skippable), threshold 9.2, path-coverage floor applies (see Phase 3). SCOPE-SKIP on a confirmed bug-fix branch = Important finding, score -1.5. State mode in executive summary triage line (`reference.md` § Executive Summary Template). Also supports `--mode=bug-fix` and `--mode=feature`. If the diff also matches the Sibling Path Trace signal (Phase 1), BugPath Verifier's dispatch payload (Phase 2 step 4) MUST additionally include the "Sibling Path Trace" excerpt from `reviewers/defect-finder.md`, since BugPath Verifier's Sibling Bug Scan dimension runs that method rather than re-deriving it.

### Phase 1: Triage

Analyze the diff and select reviewers:

| Reviewer | Focus | Activate When |
|----------|-------|---------------|
| **Defect Finder** | Correctness, edge cases, concurrency | Any code change |
| **Design Critic** | Factoring, complexity, API design | Adds/modifies classes, functions, public APIs |
| **Guardian** | Security, blast radius, backwards compat | Any code change |
| **Standards Enforcer** | Docs, test quality, observability | Always |
| **Performance Analyst** | Performance, logging | DB, loops, caching, network I/O, or >500 LOC |
| **AttackerPersona** | Credential-flow, AI-agent boundary, ident-vs-value, cookie/session, revival re-validation, CWE tagging | Any security-class signal (see signal-driven dispatch table below) or `--security` flag |
| **ShellRuntimeAuditor** | Shell/runtime portability, tool contract safety, failure-mode resilience | A shell-content signal (see signal-driven dispatch table below) — content-gated, not a path glob |
| **BugPath Verifier** | Root cause, fix coverage, sibling bugs, regression test | BugPath Mode active (see Phase 0.5) — mandatory, not skippable |
| **Monolith** (on-demand) | All dimensions | `--all` flag or manual request |

**Decision rules:** Docs-only → Standards Enforcer only. Config-only → Guardian (+ Standards Enforcer / Defect Finder only if a metric/alarm or other dispatch signal below is present). Any code → **ONE combined reviewer** carrying the Defect Finder, Guardian and Standards Enforcer lenses in a single agent, plus a mandatory `placement` dimension (one sentence + file:line: "is this the right place, or is there a simpler root cause?"; missing = incomplete review). Add separate agents only from the signal table below and the Mandatory activation rules. Measured 2026-09-20: 5 fresh agents on a 212-line diff cost 393k tokens, mostly re-reading the same files and re-running the same tests; the root-cause finding came from the placement question, which no diff signal triggers — hence mandatory.

**Signal-driven dispatch** (additive — a signal activates its reviewer(s); it never deactivates one already selected above). Scan the diff for these signals and route accordingly:

| Diff signal | Reviewer(s) | Owns |
|---|---|---|
| Metric/counter/event definition, or `.emit(`/`.inc(`/`publish(` call | Defect Finder + Standards Enforcer | Producer Trace, metric liveness, emission symmetry |
| Alarm/threshold definition, or an error branch that emits a metric or feeds an alarm | Guardian + Standards Enforcer | Alarm-feeding isolation, failure-mode differentiation, multi-provider coverage |
| External-dependency call / SDK error classifier (429, 5xx, provider error shapes) | Guardian | Multi-provider predicate coverage, monitoring blast radius |
| Field set to `null`/`0`/`false`/reset | Defect Finder + Guardian | Consumer trace, cross-cutting regressions |
| Public signature / interface / message schema / shared type change | Design Critic + Guardian | API design, backward compat |
| Loop over I/O, DB query, cache, network call, or >500 LOC | Performance Analyst | N+1, payload bloat, blocking I/O |
| File rename/move/delete | Guardian (inbound-reference focus) | Broken external consumers |
| Test-only change | Standards Enforcer (+ Defect Finder for revert-safety) | Mock fidelity, revert-safety |
| Any added `//` or `/* */` comment line embeds a ticket-tracker reference | Standards Enforcer | Ticket IDs in Comments: rot once the ticket closes, belong in commit messages and PR descriptions instead — see `reviewers/standards-enforcer.md` §3b for the full trigger and exceptions |
| Security-class signal (caller-supplied URL, dynamic SQL identifier, new MCP tool/IPC, secret read, cookie/session, `_disabled/` revival) | AttackerPersona | Credential-flow, AI-agent boundary, ident-vs-value; tags + threat-model severity multiplier |
| Shell-content signal (a shebang added/changed, a `.sh`/`.bash` file, or tool-wrapper code invoking an external CLI and interpreting its exit code/output — in any file, not just `.sh`) | ShellRuntimeAuditor | Shell/runtime portability (GNU-only flags, bash-version-gated features), tool contract safety (exit-code/flag-parsing completeness), failure-mode resilience (missing `set -euo pipefail`, silent subprocess-failure swallowing, unguarded external-binary dependencies) |
| **New user-visible feature** (new endpoint, new agent action, new UI-affecting path, new workflow branch) with **no metric or trace emit in the diff** | Standards Enforcer (OE Telemetry Gate — mandatory Critical) | Missing time-series metrics and/or distributed trace instrumentation for new behavior — ships blind; see §4a OE Telemetry Gate |
| `try { ... } catch` block in the diff (regardless of size or context) | Defect Finder -- this signal does NOT activate a new reviewer (Defect Finder already activates for any code change); it mandates specific coverage: `Catch-Swallow Fall-Through` (always) + `Finally-Block State Precondition` (only if the diff also contains a `finally` block) + `Dead Catch Verification` (before proposing any new catch) | Catch-fall-through races, finally blocks emitting against partially-completed state, unreachable defensive catches |
| Diff deletes or reroutes what was the only call site of a function/method/export, checked repo-wide (an ordinary unused-import left behind by the same refactor, with no other reference removed, is a separate minor lint concern -- NOT this pattern) | Defect Finder | Caller Removal Trace: dead code introduced by this diff (orphaned function/export) -- findings MUST use Guardian's Anti-Hallucination Gate evidence format (`reviewers/guardian.md`) |
| New property added to a type/interface/shared-state object already consumed by 2+ non-test files, OR diff touches only one file in a known parallel-path family (siblings sharing a create/update/post vs. reschedule/cancel/delete naming pattern, or sync/async twins, for the same resource) where the change supplements/supersedes a field those siblings already read | Defect Finder + Guardian | Sibling Path Trace: does every other handler for the same conceptual entity get equivalent treatment on the shared field? |

When no signal and no default activates a reviewer, skip it and say why in the triage line.

**Mandatory activation (not subject to triage exclusion):**

- **Guardian** is ALWAYS activated when changes touch: retry logic, circuit breakers, rollback behavior, deployment config, feature flags, authentication/authorization, or state machine transitions.
- **Design Critic** is ALWAYS activated in Standard Mode when changes touch: interfaces, public APIs, contracts, message schemas, shared state types, or cross-module boundaries.

**Design Critic in Bug Fix Mode (SUPPRESSED by default):** Only re-activated if diff contains API-change signals (new exports, public method signatures). State in triage line; see `reference.md` § Design Critic BugPath Logic for the detection command (if `reference.md` absent: default to SKIPPED unless `--all` flag set).
**Reduced dispatch for a cited port claim (evidence required, not self-attested):** A "verbatim/near-verbatim port of already-reviewed code" claim does NOT by itself reduce the activated reviewer set. It only reduces dispatch when the triage line pastes both the cited source (repo + commit SHA) and the literal `git diff <cited-source-sha> -- <changed files>` command plus its output — the command itself, not just the output, so the claim is independently re-runnable rather than plausible-looking. The combined reviewer is never dropped by this exemption -- it carries Guardian's lens, and a byte-identical port can still be dangerous at its new call site (different callers, different exposure). What the exemption can suppress is the signal-triggered specialist agents, for the content the citation covers. In Bug Fix Mode BugPath Verifier stays mandatory. **Carve-out:** if the cited-port diff touches concurrency-relevant constructs (locks, threads/goroutines, async/await, shared mutable state, retry/backoff timing), the combined reviewer must still run Defect Finder's interaction-path pass on it -- the exemption covers correctness/edge-case redundancy only, not concurrency. Content beyond the cited source (new tests, new logic, fixes not upstream) gets full review; the exemption never extends past what the citation covers. No evidence in the triage line means unverified — dispatch normally. (A genuine verbatim port produces a trivial/empty comparison here — if it doesn't, the "port" framing was wrong, not the review depth.) A second, distinct exemption — reduced dispatch for a small diff the reviewing agent already has current, verified context on (parallel dispatch traded for a single inline self-review pass, not a lower evidence bar) — is defined in `reference.md` § Diff-Size-Scaled Dispatch Exemption; same audit-ability requirement (state which trigger conditions were confirmed in the triage line), same rule that it reduces dispatch overhead, never dispatch substance for a risk class that's actually present.

**Overrides:** `--all` (force all), `--only=<name>`, `--skip=<name>`, `--round1-only` (skip escalation), `--security` (force AttackerPersona on), `--no-security` (force AttackerPersona off), `--mode=bug-fix` (force Bug Fix Mode), `--mode=feature` (force Standard Mode). `--skip=BugPath Verifier` is silently ignored in Bug Fix Mode (the reviewer is mandatory).

State your triage decision before dispatching:

```markdown
**Triage**: Activated: [list] | Skipped: [list] | Reason: [1-2 sentences]
```

### Phase 2: Diff + Source Context + Dispatch

Sub-agents have NO conversation context. Pass diff + source context inline.

**1. Capture diff:** `git diff --cached`, `git diff HEAD~1`, or `git diff main..HEAD`. Use `git diff main..HEAD` (or `--cached` for a single in-progress commit) whenever a full-branch-diff-dependent check applies (e.g. Ticket IDs in Comments, `reviewers/standards-enforcer.md` §3b) — `git diff HEAD~1` only shows the tip commit and silently misses anything introduced earlier in a multi-commit branch. Sub-agents have no conversation context or tool access to re-run this themselves (see above); get the scope right before dispatching.

**2. Source context for ripple analysis** (#1 missed-finding cause = reviewing diff in isolation): Fields SET/RESET/NULLED → grep READERS. Symbols DEFINED (metrics, events, enums, error codes) → grep PRODUCERS; zero producers = dead definition (mandatory finding for metric/alarm signals). Threshold comparisons → grep PRODUCERS of the compared value. Stateful code → full state type + transitions. **State machine flag context (mandatory):** when the diff touches a boolean state flag (any `is*`, `has*`, `in*`, or similar guard variable) that is SET in 2+ distinct methods in the same class, pass the FULL source of ALL methods that set or read that flag PLUS every code-path comment explaining the PURPOSE of each path. Without intent comments, sub-agents cannot detect semantic contradictions (e.g., a path whose purpose is "wait for customer speech" setting a flag that drops speech). Changed signatures → all callers. Cross-module calls → full callee body (or signature + state-mutating/throwing/early-return branches if budget-constrained). On-disk format changed (ad-hoc-parsed, no shared parser) → grep ALL consumers repo-wide incl. tests; prefer "extract a shared parser" over "update each consumer" (see reference.md Failure Modes). Symbol whose only call site the diff removes or reroutes → grep the FULL repo for remaining references; zero = dead code introduced by this diff; also pass `package.json` (`exports`/`publishConfig`/`main`) into context, since the severity ladder downgrades an exported symbol from Important to Possible when the repo is a published library (see Caller Removal Trace). New/changed field that supplements or partially supersedes an existing field on shared state → grep ALL existing readers of the field it supplements (not just the new field's own readers, which may not exist yet) → verify every such reader was updated, or confirm why it doesn't need to be (see Sibling Path Trace).

**3. Inbound reference scan** (mandatory when diff renames, moves, or deletes files):

```bash
git diff --diff-filter=RD --name-only main..HEAD   # old paths
grep -rn "old-filename" . --include="*.md" --include="*.ts" --include="*.sh"  # scan ENTIRE repo
```

**MUST scan outside the changed directory.** The #1 failure mode: scoping grep to the refactored directory, missing sibling modules that reference old paths. Hits outside the diff are **mandatory CRITICAL findings** — broken consumers the author didn't update. Include grep results in every reviewer's context.

**4. Dispatch ALL activated reviewers simultaneously** via `sub-agent-code-reviewer` (Augment) or `Task()` (Claude). Each gets: reviewer prompt + full diff + source context + inbound reference scan results. In BugPath Mode with the Sibling Path Trace signal also matched (Phase 1), this is the step where BugPath Verifier's payload must actually get the "Sibling Path Trace" excerpt promised in Phase 0.5 -- not just noted as a requirement earlier.

### Phase 3: Aggregate

After all reviewers return:

1. Sort findings: **Critical → Important → Minor → Possible**, then by file path
2. Prefix each with `[Reviewer Name]`
3. **BugPath Verifier SCOPE-SKIP** (BugPath Mode only): if the BugPath Verifier reports SCOPE-SKIP, add an Important finding "BugPath Verifier SCOPE-SKIP on confirmed bug-fix branch — manual path-coverage review required" and deduct 1.5 from the score.
4. **Convergence**: same location from 2+ reviewers — keep both; True convergence (different reasoning paths) → promote to **≥ Important** (never demote a Critical); Echo convergence (same evidence/phrasing) → retain original severity.
5. Clean dimensions need same `evidence` block as findings; missing evidence on any clean dimension causes the verifier to cap the overall run score at 7.0.
6. **Severity**: Critical=broken now; Important=breaks under conditions; Minor=standards gap. Elevate to Important when operator-visible signal is wrong/missing. Reclassify process gaps downward. See `reference.md` § Severity Definitions.
7. **Triple-filter** each Important/Critical on CX impact, complexity, testability: **Implement** (all 3 pass) — propose exact code change + enumerate `collateral_test_changes` (required, may be empty-list: grep `it(`/`test(`/`describe(` in diff's test files for tests whose description references the behavior being removed; only include if assertion CLEARLY validates that specific behavior). **Defer** — document for future work. **Reject** — fix adds more complexity than it removes. Preserve Regressions Risked + Durable Check per Implement finding.

**Tightening**: >10 findings → suppress Minors from body (count in summary; state "Tightening applied: [N] Minor findings suppressed"). **Score**: `10.0 − 2.5×C − 1.5×I − 0.25×M − (durable<50%?0.5:0)`, floor 0.0. Extract threshold from invocation (e.g. `/sp-cr-battery 8.5` → 8.5; default 7.0, or 9.2 in BugPath Mode). Verdict is anything other than PASS or PASS_WITH_NITS (below-threshold score, OR a Reject-classified Critical at or above threshold) → skip the sentinel write step in Phase 6 (still write the JSON envelope to `.cr-battery-runs/`). BugPath path-coverage floor: INSUFFICIENT→cap 6.5, PARTIAL→8.0, FULL→none. Metrics: durable ≥50%, convergent count, unresolved Critical=0.

**Report format**: Executive Summary (see `reference.md` § Executive Summary Template) → Header → Critical → Important → Minor → Possible → Clean Dimensions → Action Classification → Durable Checks → Summary (`Findings: [N]C/[N]I/[N]M ([N] suppressed)/[N]P | durable=[N]%, convergent=[N], unresolved-critical=[N]`).

### Phase 4: Escalation (Round 2)

If ANY trigger fires after Round 1, re-dispatch a focused reviewer:

| Trigger | Re-run | Why |
|---------|--------|-----|
| >2 state/flag findings OR diff touches a boolean state flag (`is*`, `has*`, `in*`) set in 2+ methods in the same class | Defect Finder (interaction-path focus) — apply Phase 2 state machine flag context rules for the dispatch (full setter/reader source + purpose comments) | Semantic contradiction risk: flag purpose vs. flag effect may contradict across paths |
| >3 test quality issues | Standards Enforcer (mock-focused) | Shared mock infrastructure |
| >50 lines removed or functions deleted | Guardian (deletion focus) | Callers may depend on removed behavior |
| Files renamed/moved/deleted without inbound scan | Guardian (inbound-reference focus) | Broken consumers outside the changed directory |
| "Pre-existing" issues flagged | Defect Finder (lifecycle focus) | Deeper structural gaps |
| Diff adds/changes a metric or alarm definition, or an error-handling branch that emits a metric or feeds an alarm | Standards Enforcer (observability-completeness focus) | Dead definitions, Success/Failure asymmetry, undifferentiated failure modes |
| Diff adds new user-visible functionality with zero metric/trace emits (OE Telemetry Gate) | Standards Enforcer (OE focus — mandatory Critical) | Feature ships with no time-series metrics or trace instrumentation; cannot be operated or alarmed on |

**Collect every trigger that fires and dispatch all of them together as ONE round** (same "simultaneous" dispatch discipline as Phase 2 step 4) — e.g. if both the interaction-path and inbound-reference triggers fire, that is still Round 2, one dispatch, not two sequential round-trips. Firing triggers one-at-a-time adds pure serial latency with no rigor gain: dispatching N triggered re-checks together costs the same wall-clock as dispatching one. Continue the Round 1 reviewer where possible, sending the diff slice + trigger signal and stating its earlier context is superseded for changed files; a fresh agent needs the complete Round 1 findings, not a summary. Append under `### Round 2 Findings`. Skip if `--round1-only`, all clean, or diff <20 lines.

### Phase 5: Convergence

**STOP** when: unresolved Critical = 0, last 2 passes <20% new high-sev, durable ≥50%. **CONTINUE** if escalation trigger fires or Critical remains. **ESCALATE TO HUMAN** after 3 passes.

### Correlated-Failure Detection

After synthesis: (1) evidence overlap: ≥3 reviewers cite same file+line → flag and expand scope; (2) phrasing similarity: near-identical rationale across findings → flag and re-examine from different entry; (3) clean sweep: all-zero findings → verify reviewers examined different slices. Flags trigger scope expansion, not verdict changes. See `reference.md` § Correlated-Failure Detection.

### Phase 6: Finalize Verdict + Write Sentinel

**Prerequisite:** Correlated-Failure Detection has completed and no re-examination was triggered.

**Preserve the run (before sentinel write):**

1. Determine the verdict from this MR's score vs threshold: PASS if score >= threshold (no unresolved nits); PASS_WITH_NITS if at or above threshold but Minor nits remain; PASS_WITH_FIXES if below threshold, every Critical finding is Implement-classified (fixable path exists — vacuously true if there are no Critical findings), and every Important finding is Implement- or Defer-classified with a documented rationale in the report (llm-skill-review dogfood review, 2026-07-17: a Critical is NEVER Defer-classified, with or without rationale — that always forces REJECT; a follow-up pass the same day found the first fix still left exactly this Critical+Defer combination unmapped); REJECT if any Critical is Reject- or Defer-classified, any Important Defer has no documented rationale, or there are unresolvable blockers.
2. Write a JSON envelope to `.cr-battery-runs/<HEAD-sha>.json` with `tools/review-envelope.py` (`init --kind battery`, `add-clean`/`add-finding`, `resolve`, `set`, `check`): it runs each evidence command as the claim is written and refuses one whose expectation fails. Schema: `reference.md` § Run Envelope Schema. Every finding AND clean-dimension verdict must carry an `evidence` block; `verifiable: false` caps at 7.0. Expectation types: `reference.md` § Verifier Details. `tools/run-battery.sh` refuses sentinel write if JSON missing in Bug Fix Mode; graceful degrade in Standard Mode. **Note:** `collateral_test_changes` is not yet captured in the run envelope schema — it appears only in the human-readable report. Tracked in `candidates/candidate-009.yaml` for future schema update.

If final verdict is `PASS` or `PASS_WITH_NITS` (all nits resolved):

```bash
# tools/run-battery.sh is the ONLY permitted way to write .code-review-cleared -- never echo it directly.
tools/run-battery.sh --verdict PASS --min-score "<threshold>"         # no unresolved nits
tools/run-battery.sh --verdict PASS_WITH_NITS --min-score "<threshold>" # Minor nits remain unresolved
```

**Timing:** Pre-commit battery → sentinel stales after commit; re-run before push. `REJECT` or `PASS_WITH_FIXES`: do NOT write sentinel — fix Critical/Important, re-dispatch, re-run.

### Gap Analysis + Error Handling
Monolith found something no specialist found → candidate pattern → `candidates/`. Reviewer fails → note, don't retry. Diff >3000 lines → warn, suggest chunks, keeping structurally-related files (Caller Removal Trace or Sibling Path Trace candidates) in the same chunk or cross-passing their excerpts so neither chunk's reviewer works blind to the other file -- this bounds what one dispatch reads, not what Phase 2's full-repo grep obligation searches; a chunked sub-battery still greps the whole repo for out-of-diff references. Empty diff → skip.

## Failure Modes & Anti-Patterns

See `reference.md` for the 5 standard failure modes (no-findings, FPs-from-isolation, convergence-stuck, monolith-vs-specialist, duplicated-format-parsing) and the 5 anti-patterns (all-agree, duplicates, fatigue, missing-context, over-scoping), with a companion table each (detection/correction for anti-patterns, a fix column for failure modes). Moved to the companion file to keep this skill under the per-skill line budget.

## Companion Skills

- **progressive-code-review-gate**: Primary consumer (dispatches this battery pre-commit) · **providing-code-review**: Engineering rigor checklist (informs reviewer focus)
- **inter-agent-review-protocol**: File-protocol review (alternative dispatch method) · **micro-harsh-review**: Per-batch review

Attribution

bordenetbordenet
View sourceSee grades on GitHubMore from bordenet →
SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments (0)

No comments yet. Be the first to comment!

SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Related Skills

Caveman

Terse caveman voice: answer first, fluff gone, every technical fact kept. Use for /caveman, "caveman mode", "talk like caveman", "be brief", "less tokens". Stays on until "stop caveman" or "normal mode".

1100021 votes

Hyperplan

Adversarial multi-agent planning skill. Self-orchestrates 5 hostile category members (unspecified-low, unspecified-high, deep, ultrabrain, artistry) via team-mode for ruthless cross-critique debate, distills only the defensible insights, then MANDATORILY hands the distilled insight bundle to the `plan` agent for executable plan formalization. Use when planning needs maximum rigor and surfacing of weak assumptions, blind spots, and over-engineering. Triggers: 'hyperplan', 'hpp', '/hyperplan', ...

698621 votes

Writing Skills

Create and manage Claude Code skills in HASH repository following Anthropic best practices. Use when creating new skills, modifying skill-rules.json, understanding trigger patterns, working with hooks, debugging skill activation, or implementing progressive disclosure. Covers skill structure, YAML frontmatter, trigger types (keywords, intent patterns), UserPromptSubmit hook, and the 500-line rule. Includes validation and debugging with SKILL_DEBUG. Examples include rust-error-stack, cargo-dep...

3931 votes

Mcp Code Execution

Routes multi-tool workflows through MCP servers for large datasets and pipelines. Use when Bash tool overhead is limiting throughput on data-heavy tasks.

3421 votes

catchup

Recovers the conversation and failed tool calls of a previous Codex, Amp, Claude Code, Antigravity, Cline, Copilot CLI, Cursor, DeepSeek Harness, Grok Build, Kimi, OpenCode, Pi Agent, or ZCode session. Use when the user says "catch up", "what did the last session do", "get me up to speed", "I switched agents", asks to recover/summarize a previous session before continuing, or asks to diagnose or report a catchup failure. Do NOT use for the current conversation, git history, or any non-agent log.

741 votes
View all in ai-agents →