Skip to content
Back to skills

Test Audit

ASecurity

Use when writing, reviewing, or pruning tests. Three modes, one value bar: a four-question gate before any new test is written, a focused audit of one scope, and a campaign over a whole repo or subsystem that removes the least useful tests under a coverage guard. Every delete, merge, or repair carries written evidence and a caught mutation. Triggered by "audit the tests", "prune the test suite", "remove low-value tests", "test diet", "is this test worth adding", /test-audit.

  • 53 stars
  • 0 votes
  • 0 copies
  • 0 views
  • Added October 2, 2026
ai-agentsgobashtestinggitapidatabasesecurity

Works with

  • api

Security analysis

A100/100

Pro scans all 3 files and shows the line behind each finding

Scanned October 2, 2026

npx -y skills add Kanevry/session-orchestrator --skill test-audit --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Test Audit?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Test Audit
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/kanevry-test-audit-1edb89e9/badge)](https://www.skillsdirectory.com/skills/kanevry-test-audit-1edb89e9)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: test-audit
description: >
  Use when writing, reviewing, or pruning tests. Three modes, one value bar: a four-question gate
  before any new test is written, a focused audit of one scope, and a campaign over a whole repo or
  subsystem that removes the least useful tests under a coverage guard. Every delete, merge, or
  repair carries written evidence and a caught mutation. Triggered by "audit the tests", "prune the
  test suite", "remove low-value tests", "test diet", "is this test worth adding", /test-audit.
user-invocable: true
argument-hint: "[gate|audit|campaign] [scope] [--goal \"<goal>\"]"
model: inherit
---

# Test Audit

A test earns its place by catching a regression nothing else catches. This skill applies that bar at three depths. It complements `.claude/rules/test-value.md` (TV-001..TV-005) and the repo's test-quality / test-hygiene rules; it replaces none of them.

Rule references below (`test-value.md`, `receiving-review.md`, `parallel-sessions.md`, `bash-harness-pitfalls.md`) name the repo's `.claude/rules/`. A consumer repo often lacks them; then read them from the plugin's `rules/always-on/`.

## Invocation

Invoked as `/test-audit [gate|audit|campaign] [scope] [--goal "<goal>"]` with arguments: **$ARGUMENTS**.

- `gate` — run the write gate on the test about to be written. Default when no mode is given and the conversation is about to add a test.
- `audit <scope>` — one focused pass over a path, glob, or module. Default when a scope is given without a mode.
- `campaign <scope>` — a whole repo or subsystem, split into lanes. Only on the explicit word `campaign`; read [references/campaign.md](references/campaign.md) before starting.
- A missing scope for `audit` or `campaign` is asked for once via `AskUserQuestion`, with the largest test directories as options.

## Set an ambitious goal first (audit and campaign)

State the goal in the first output and repeat it in the report. With no `--goal`, use:

> Remove at least 20 % of the least useful tests, measured in test lines. Total coverage stays within 2 percentage points of M0, and no production file drops below its own M0 coverage.

Test lines count only for files the CI actually runs; tracked test files outside that scope are reported on their own line, as dead or as live outside the include, and never counted toward the goal ([references/campaign.md](references/campaign.md) § Measurements).

The goal exists because a bare "clean up" stops far too early. It never licenses a deletion without evidence. What counts toward it, and what is reported beside it (same-file reshaping, mock ballast, F/N growth, the O sum), is fixed in [references/campaign.md](references/campaign.md) § Goal accounting. A goal that is only reachable by breaking the evidence rules below is not reached: the report says which rule blocked it and where the audit stopped. An audit that ends with zero changes is valid when the ledger shows why.

## Mode 1: write gate

Before adding any test, answer all four. A missing answer means the test is not written yet.

1. Which observable behaviour, invariant, or independent contract does it protect?
2. Which credible regression turns it red?
3. Why does no existing test catch that regression? Grep first (TV-004); extend a table case or shared fixture instead of adding a near-duplicate.
4. Does it need a production seam (export, flag, hook, injection parameter) that no production caller uses? Then no: move the test to the real boundary. An export or parameter with one production caller is API when the package's public entry (`exports`, index, another repo) reaches it, and a seam otherwise.

Then check it against the junk patterns below; a match fails the gate unless the keep list names the contract it alone guards. A regression test must be shown red on the pre-fix code, for the intended reason, before the fix lands; one that never failed proves the mock, not the fix. One regression at the owning boundary covers the bug; do not replay it at every layer. `no-tests-needed: <reason>` is a success outcome (TV-001).

## The value bar

Judge each test by what its assertions can catch, never by its name.

**Junk patterns** (candidates for D, C, or F):

- no assertion, or an assertion that cannot fail (self-comparison, identity copy);
- copied inventories, export lists, manifests, or fixtures that restate the source;
- a test-local copy of another repo's schema or contract, checked against test-local payloads: it proves the copy, not the contract; the keeper is the synced artefact or the other repo's test;
- exact source, import, or string greps where a behavioural check exists;
- the expected value is produced by the helper under test;
- the mock implements the behaviour being asserted, or one mock stands in for different APIs;
- a negative control that passes for an unrelated reason (a different guard rejects first, or the production path never reaches the rejection); for a security check, one that only hits the early exit (wrong length, wrong prefix) and never the comparison itself is usually F;
- the name promises more than the assertion checks;
- an assertion that depends on the current date (a fixture expiring next quarter turns it red on its own; F with a fixed clock, and urgent);
- prose or structure pinned in a `.md` file (TV-002c), unless the project instruction file makes that structure a rule (then R, citing it);
- the test is the only caller of dead production code: D only together with the removal of that code (its own `code-implementer` commit, owner decision when the code is public API), otherwise O.

**Keep list** (the counterweight to TV-002): a test that is the independent proof for a public API, protocol, a config value production reads (a config field whose only reader is the test is a seam, not a contract), migration, storage format, security control, release step, or data-protection rule stays. So do observable call ordering and regressions with a credible failure mode. Static or slow is not a reason to delete. When in doubt, mark R.

**Accidental coverage** is not a proof: a line that runs only because a mock is incomplete (a missing method sends the code into its catch branch) or because a timer is left dangling after the test. Deleting its duplicate loses no contract but trips the coverage backstop; answer that with a contract test (fake timers, an explicit failing mock), never by restoring the accidental one.

## Ledger marks

Every test declaration in scope gets exactly one mark and one evidence line. A parametrised table is one declaration unless its rows need different marks.

| Mark | Meaning | Required note |
|---|---|---|
| R | keep | the contract and the bug it catches |
| F | keep the contract, repair the assertion | what the assertion misses today, and the mutation the repair must catch |
| C | consolidate into a keeper | the keeper test that absorbs it |
| D | delete | the remaining proof, or why no contract exists |
| N | new test | only after passing the write gate |
| B | red on the baseline (from the P1 / M0 pass/fail list only) | possible product bug: never deleted, becomes an issue |
| O | owner decision pending: keep | the open question and the candidate's line count; listed in the MR under open decisions, never folded into R |

Suspected product bugs found while reading, with no test behind them, go into a separate _Issues_ section of the ledger (`file:line`, expected vs. actual). They become issues, never red tests in the MR.

## Evidence before any C or D

All eight fields, written into the ledger before the edit. A missing field means no deletion.

1. Test location (`file:line`) and exact name.
2. The failure it can actually detect.
3. Non-test callers of the code it covers.
4. The stronger proof that remains (the keeper), or why none is needed.
5. History: `git log -S "<identifier>" -- <file>` and the reason the test or seam exists.
6. What the deletion unlocks (a test-only export, a wrapper, a dead path).
7. Risk and the focused command that re-checks it.
8. The proposed contract mutation (id, `src/file:line`, and the diff as a patch file) under which the keeper must go red; falsification executes it. A D that removes no contract (dead code, a duplicate proven elsewhere, a test-local copy) says so here and becomes a `NOCONTRACT` manifest row instead ([references/campaign.md](references/campaign.md) § Mutation manifest).

## Falsification

For every contract that a C, D, or F relies on, make one targeted mutation in the production code: flip the condition, drop the field, change the byte the contract names. Several marks may share one mutation when they name the same keeper and contract; the ledger maps every mark to at least one mutation id. Replacing a whole function body with `throw` kills every caller and proves nothing; use it at most as a pre-filter. The keeper (for C/D) or the repaired test (for F) must turn red under the mutation.

- Apply the mutation as a patch (`git apply`), run the one test, reverse it (`git apply -R`), then prove the restore byte for byte: `git diff --exit-code -- <prod-file>`, or a checksum taken before the mutation when the file already carried edits. A leftover diff aborts the audit. [references/mutate.sh](references/mutate.sh) runs a manifest of such patches ([references/campaign.md](references/campaign.md) § Mutation manifest). Prove the red with the runner's failure summary (vitest: `Tests N failed`), never with the exit code alone, and keep the output files.
- Never mutate a checkout another session or agent is using, and never edit while the test runner is live in that checkout. The binding run is the coordinator's, in an offload worktree: `offload run <repo> -H <host> --job ta-<repo>-<purpose> -- <command>`, reused with `--no-sync` (hosts come from `remote-hosts:` in Session Config; see the `remote-offload` skill; path and expansion pitfalls in [references/campaign.md](references/campaign.md) § Measurements). An agent's own pre-check runs only in a private copy outside the repo, under [references/campaign.md](references/campaign.md) § Pre-checks, and never counts as proof.
- Log each mutation: `file:line`, the mutation diff, the test that went red, the restore proof.
- Per-file coverage is the mechanical backstop, compared on all four metrics (lines, branches, functions, statements); lines alone miss a lost callback. When any of them drops for a production file, locate the line, branch, or function in the detailed report (vitest `--coverage.reporter=json`) and restore the contract in the keeper with a caught mutation. Restore the deleted test verbatim only when it was the genuine proof; accidental coverage (§ The value bar) gets a contract test instead.

Bash harnesses are falsified before their tests are judged: feed a known-broken input and require a non-zero exit and a FAIL in the written artefact, under `bash`, not zsh (`.claude/rules/bash-harness-pitfalls.md`).

## Mode 2: audit

1. Read the project instruction file and the test rules it points at. Check the coverage provider and the CI scope first, then record the M0 numbers for the scope at a pinned SHA (recipes in [references/campaign.md](references/campaign.md) § P1 and § Measurements).
2. Read every test in scope in full, plus the production owner, its entry point, callers, and history. For more than a handful of files, dispatch `qa-strategist` read-only to draft the ledger; it writes `/tmp/<audit>/ledger-<lane>.md` itself and answers with the count line, and the coordinator assembles the ledger file. A scope above one lane (about 2,200 test lines or 150 declarations, [references/campaign.md](references/campaign.md) § P2) needs campaign mode: say so and ask for the word `campaign` before step 3.
3. Write the ledger with marks and evidence. Prefer a few well-proven candidates over a long speculative list.
4. Keepers first, in one uncommitted pass: `test-writer` repairs every F, absorbs every C into its keeper, and deletes the C and D. Agents never commit (PSA-007 in `.claude/rules/parallel-sessions.md`).
5. Falsify every C, D, and F against that tree; commit only when every run row is `CAUGHT`. Removing a test-only production seam is a separate `code-implementer` task and commit.
6. Independent preservation review, read-only (`session-reviewer`, or `pr-review-toolkit:pr-test-analyzer` when installed): contracts that lost their only proof, and new assertions that cannot fail. Every restored contract needs a caught mutation.
7. Measure M1 with the same commands, then report.

## Mode 3: campaign

Phases P1 to P9 with lanes, a separate layer-plan pass, an owner checkpoint after P4, the sub-campaign threshold, the pre-check rules, and the full measurement set live in [references/campaign.md](references/campaign.md). Read it completely before P1.

## Never without the owner

Ask via `AskUserQuestion`, and do not work around a refusal. A headless run without that tool takes the answer the brief or this list already gives and logs it; anything else becomes an O mark plus an entry in the run's parking file (the report's _Parked_ section when the brief names none), never a workaround:

- changing coverage thresholds, include/exclude lists, or runner config (never `autoUpdate`);
- deleting or skipping e2e, Playwright, or real-database suites;
- editing CI config or git hooks;
- updating snapshots with `-u`;
- adding any dependency, a mutation-testing tool or a missing coverage provider included (raised in P1, with the three options in [references/campaign.md](references/campaign.md) § P1);
- fixing a product bug found on the way (it becomes an issue);
- removing a seam whose removal changes a public contract (`stop-and-escalate`, RCR-007 in `.claude/rules/receiving-review.md`);
- merging.

## Stop and success

Stop when: a change you did not make, or a lock, appears in scope (PSA-002); a restore leaves a diff; the baseline does not reproduce; two review cycles pass without fewer open findings (RCR-008: land the smallest safe subset); a lane exceeds its budget.

Done when: the gate is green at the MR SHA; flakes and per-file coverage are no worse than M0; every C, D, and F has evidence and a caught mutation (a D without a contract may instead have a `NOCONTRACT` row with ledger evidence); the production diff against `main` outside the test globs and `docs/audits/` is empty, seam commits and owner-approved changes excepted; the preservation review has no open gap; the goal is reached, or the report names why it was not; the MR pipeline has finished and its result stands in the report.

## Output

- `docs/audits/<date>-test-audit.md` (report) and `docs/audits/<date>-test-audit-ledger.md` (ledger) in the target repo.
- A draft MR titled `test(audit): <scope> Test-Audit <date>` (subject case per the repo's commitlint config) carrying: M0/M1 table, R/F/C/D/N/B/O counts per lane, the keeper per contract, the mutation table, kept false alarms and why, production and test lines counted separately, the zero-production-diff proof, owner decisions taken and open (every O), and the pipeline result.
- Commits by the coordinator: one per lane, owner lanes before the cross-cutting lane, seam removal separate, the audit documents separate (commitlint and clone depth: [references/campaign.md](references/campaign.md) § P6, § P9).
- Follow-up issues: every B, every entry of the ledger's _Issues_ section, every flake, every seam whose removal changes a contract, every N gap not implemented.

Inspired by openclaw/openclaw `.agents/skills/test-audit` @ 80930af (MIT, Copyright (c) 2026 OpenClaw Foundation); rewritten for session-orchestrator.

Files in this skill

  • SKILL.md12.9 KB
  • references/campaign.md12.5 KB
  • references/mutate.sh2.6 KB

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…