Audit, prune, or write tests in Agent Knock Knock. Use for periodic test-value reviews and for test changes where duplicate or implementation-coupled coverage may be introduced.
Installs into .claude/skills of the current project.
Are you the author of Test Audit?
Add the live security badge to your README. It updates with every re-scan.
[](https://www.skillsdirectory.com/skills/scotthuang-test-audit)
---
name: test-audit
description: Audit, prune, or write tests in Agent Knock Knock. Use for periodic test-value reviews and for test changes where duplicate or implementation-coupled coverage may be introduced.
---
# Test audit
This is the repository adaptation of OpenClaw's [test-audit skill](https://github.com/openclaw/openclaw/blob/80930af448ebabc84174146b56bc106d37fab3b4/.agents/skills/test-audit/SKILL.md) and [campaign guide](https://github.com/openclaw/openclaw/blob/80930af448ebabc84174146b56bc106d37fab3b4/.agents/skills/test-audit/CAMPAIGN.md), reviewed on 2026-09-25. See [the OpenClaw MIT notice](LICENSE.openclaw). Compare the pinned source with upstream when refreshing this skill; retain this repository's test policy and commands.
Use an **authoring gate** for new or changed tests, a **focused audit** for a bounded owner, or a **campaign** for a complete subsystem/repository inventory. Read [CAMPAIGN.md](CAMPAIGN.md) before a campaign. Optimize for independently protected behavior, not deleted line count.
## Repository boundary
Read root and scoped `AGENTS.md`, [testing documentation](../../../docs/testing.md), and the current `package.json` scripts before execution. `test/test-tiers.json` is the authoritative root test inventory; include connector tests, UI tests, scripts, fixtures, and dynamically generated cases when they belong to the chosen scope. Check `config/public-contract-witnesses.json` and architecture/evidence validators before changing an apparent source-inspection test or its named witness.
During ordinary development, review, and audit, run only the fast test tier: `npm run test:fast`. For touched connectors, their fast scripts (`npm run pi:test:fast`, `npm run deepseek:test:fast`) may add scoped evidence. `npm test`, `test:full`, `test:integration`, `test:affected`, and release tiers are not audit commands: `test:affected` can select the full tier. Full/release tests belong only to the immediate gate for an actual npm or ClawHub publication. Typecheck, builds, `npm run validate:architecture`, `npm run validate:refactor-evidence`, and `npm run skill:check` are non-test checks and may run when relevant. Report integration/full proof as unrun under policy rather than implying it passed.
## Authoring gate
Before adding or materially changing a test, identify:
1. The observable behavior or independent contract it protects.
2. A credible regression that makes it fail for the intended reason.
3. The strongest existing owner test and the distinct risk this case adds. Prefer extending a table or fixture over replaying one contract at several layers.
4. Any production export, flag, wrapper, or injection hook required only by the test. Prefer the real owner boundary instead.
For a bug regression, demonstrate the failure before the fix and the pass after it when the permitted fast tier can exercise that boundary. If only a release-tier test reaches it, retain the case when justified and state that its execution awaits the authorized publication gate.
## Audit judgment
Treat these as leads, never automatic deletion rules:
- Assertion-free probes, self-comparisons, or expected values generated by the function under test.
- Copied manifests, fixtures, export lists, source-text greps, and call-shape assertions that break under behavior-preserving refactoring.
- Repeated tests of one contract across layers or connector-local replays of a shared helper.
- Mocks that supply the behavior, receipt, persistence, or ordering that the asserted production path should produce.
- Negative cases that pass through an unrelated guard, capability checks that only repeat flags, and test names that claim more than the assertions prove.
- Tests that solely preserve a test-only production seam, or production code whose only callers are tests.
Retain independent public API, host/plugin, protocol, schema, migration, Store, terminal, security, default, compatibility, package, release, and architecture contracts. Observable ordering and credible regression guards also remain valuable. Source inspection can be the cheapest independent guard for an exact public key, byte, path, or required architecture constraint; confirm it would survive an identifier-only refactor. Slow or static alone is not grounds for deletion. A failing baseline test may expose a product defect.
Before classifying a candidate, read its full test (including table rows), production owner, entry points and non-test callers, sibling implementations, overlapping coverage, test-tier routing, and relevant history. Use the assertions to judge the test, not its name. Discovery is read-only until the evidence below is recorded.
## Candidate record
Record all seven fields before deleting or consolidating a test or assertion block:
1. Exact test name and location.
2. What failure its assertions can actually detect.
3. Non-test callers of any covered production or support seam.
4. Stronger remaining owner-boundary proof, or why no proof is needed.
5. Relevant history and why the test or seam was added.
6. Production or test-support deletion made possible.
7. Risk and the permitted validation command or reason execution is deferred.
Mark each case `R` (retain), `F` (repair), `C` (consolidate into a named keeper), or `D` (delete with a named keeper or evidence that no contract exists). When one case mixes valuable and redundant assertions, mark the declaration `R` and annotate the specific assertion blocks with their `F`/`C`/`D` action and keeper; do not delete the entire case. For a focused audit, report only high-confidence candidates; for a campaign, account for every case as described in [CAMPAIGN.md](CAMPAIGN.md).
## Cutover and handoff
Edit one coherent owner boundary at a time. Move unique assertions into the keeper before removing duplicated layers. Remove obsolete test-only seams with their last callers. Update `test/test-tiers.json`, public-contract witnesses, and any generated/evidence inventories only when their actual ownership changes; do not weaken a gate to make deletion pass. Preserve user work in a dirty checkout by isolating edits.
After the final edit, run the root fast tier and touched connector fast tiers, plus relevant non-test validators and `git diff --check`. Do not edit source or tests during an active test run in the same checkout. Independently compare deleted assertions with keepers; where feasible, make a deliberate owner mutation and prove a fast-tier keeper fails, then restore the source byte-for-byte. If the owner is integration-only, document the unexecuted proof gap rather than bypassing the test policy.
Report the contracts kept, repaired, consolidated, and deleted; any false positives; before/after test and production line counts separately; checks actually run; and deferred release-tier evidence. Commit or push only within the task's authorization.