Run the gauntlet, implement verifiable user stories, or set up agent quality gates. A staged pipeline specifies Gherkin, implements to green tests, cleans to a per-function CRAP threshold, and hardens through mutation testing.
Installs into .claude/skills of the current project.
Are you the author of Agent Gauntlet?
Add the live security badge to your README. It updates with every re-scan.
[](https://www.skillsdirectory.com/skills/n0an-agent-gauntlet)
---
name: agent-gauntlet
description: Run the gauntlet, implement verifiable user stories, or set up agent quality gates. A staged pipeline specifies Gherkin, implements to green tests, cleans to a per-function CRAP threshold, and hardens through mutation testing.
license: MIT
metadata:
author: Anton Novoselov
version: "0.4"
---
# The Agent Gauntlet
Uncle Bob's multi-agent methodology, as a portable skill. Core idea: **don't stuff quality rules into prompts - wrap the agent in deterministic gates.** Agents treat written rules as vague suggestions and forget the middle of long instructions; tools don't decay. Put the agent in a loop against a checker until the numbers pass.
## The pipeline
```
story ──▶ specifier ──▶ coder ──▶ cleaner ──▶ hardener ──▶ (qa: phase 2)
Gherkin + tests + CRAP gate mutation
QA proc impl passes score clean
```
Each stage runs in a **fresh context** and **its own git worktree**. Only disk artifacts cross stage boundaries: `features/<slug>.feature` and `features/<slug>.qa.md` are the contract; commits are the code; `.gauntlet/runs/<slug>/handoffs/` carries reports. This prevents context contamination, preserves the developer's checkout, and makes each stage's changes diffable.
## Stage contracts
Run stages as separate sessions using their system prompts in `stages/{specifier,coder,cleaner,hardener,qa}.md`; without subagents, read and follow each directly. Summary:
1. **specifier** - story text → `features/<slug>.feature` (3-7 Given/When/Then scenarios, user-observable only, no implementation terms) + `features/<slug>.qa.md` (numbered manual QA steps). Never writes code.
2. **coder** - feature file → unit tests + implementation until green. May not edit the feature file; if the spec is unimplementable, it stops and reports. No refactoring of surrounding code.
3. **cleaner** - changed files → tests or refactors until every function is under the CRAP threshold. Behavior frozen: structure changes only, tests stay green, assertions never weakened.
4. **hardener** - changed modules → mutation testing until zero unjustified survivors. Never weakens production code. Requires current raw tool evidence plus regular tests; narrow waivers name exact mutant IDs, reasons and reviewers. Follow `references/mutation-evidence.md`.
5. **qa** (phase 2, optional) - QA procedure → executable UI check with screenshot evidence, PASS/FAIL per scenario.
Pipeline rules: stages run sequentially; a failing stage stops the pipeline (no skipping); stages commit only in their own worktree, nothing is merged or pushed - the run's only output is the branch `gauntlet/<slug>`, which the human reviews.
## The handoff protocol
`scripts/gauntlet/gauntlet.sh` is the runtime (Bash 3.2 + git, Python 3.8+ for mutation evidence, no daemon). Full contract in `references/handoff-protocol.md`; the short version:
```bash
scripts/gauntlet/gauntlet.sh start <slug> --story story.md [--qa]
scripts/gauntlet/gauntlet.sh stage <slug> coder
scripts/gauntlet/gauntlet.sh handoff <slug> coder --gate "swift test"
scripts/gauntlet/gauntlet.sh status <slug>
scripts/gauntlet/gauntlet.sh finish <slug> [--purge]
```
`handoff` refuses dirty worktrees, missing `By <stage>.` bylines, non-descendant HEADs, coder edits to `features/`, missing reports, red gates, and gate-induced tree changes. Hardener always needs regular tests and valid commit/base/scope-bound mutation evidence. Logs and captures survive `finish`. The first valid call returns `AUDIT_REQUIRED` (4); recheck the checklist, then repeat unchanged to queue and advance the branch. Changes to commit, report, gate or evidence restart audit. Other exits: 0 success, 1 refused, 2 setup, 3 `NO_TASK`.
Stages begin with `stage`, use only its printed payload, work/commit in `WORKTREE`, write `.gauntlet/report.md`, and end with `handoff`. The orchestrator launches fresh stages named in `NEXT:` and resumes through `status`.
## The deterministic gates
| Gate | Tool | Stage | Threshold |
|---|---|---|---|
| Tests green | `swift test` / project test cmd | coder, hardener | 100% pass |
| CRAP per function | lizard + coverage export, joined by `crap.py` via `crap-gate.sh` | cleaner | 6 (`GAUNTLET_CRAP_THRESHOLD`) |
| Module line coverage | llvm-cov totals, SPM mode only | cleaner | floor 70% (`GAUNTLET_COV_FLOOR`, 0 disables) |
| Structure (optional) | SwiftLint `swiftlint-gauntlet.yml`, when installed | cleaner | body 100 lines, nesting 3 |
| Mutation evidence | `mutation-evidence.py`, raw Muter or standard report JSON | hardener | nonzero results, exact commit/scope, zero unjustified survivors |
CRAP = `complexity^2 * (1 - coverage)^3 + complexity`, per function. At full coverage a function scores exactly its complexity, so a threshold of 6 means "at most six paths, all of them tested". Uncoverage is cubed: complexity 12 at 0% scores 156. That is what makes it a change-risk gate rather than two disconnected numbers - a module-level coverage floor lets an untested complex function hide behind well-covered neighbours; the per-function score does not.
Thresholds are Uncle Bob's agent-calibrated values: he held humans to CRAP 4 and gives agents 6 (perfect short-term memory handles more complexity), and is trying 8. Adjust *thresholds* for agents, keep *values*; don't impose human *disciplines* (no strict TDD choreography).
### Other stacks
Complexity is lizard in every stack (27 languages, pure Python, auto-installed into `.crap/venv`). Coverage is whatever the repo already exports; hand it to the gate:
```bash
scripts/gauntlet/crap-gate.sh --lcov coverage/lcov.info src # c8/nyc/jest/vitest, coverage.py lcov, gcov
scripts/gauntlet/crap-gate.sh --cobertura build/coverage.xml src # JaCoCo, .NET, coverage.py xml
scripts/gauntlet/crap-gate.sh --xcresult .crap/cov.xcresult Sources # iOS/macOS APP project
```
Mutation validation is language-neutral: Muter v16 capture for Swift, or `run-report` for a committed command producing Mutation Testing Report v1/v2 JSON. A StrykerJS runner is bundled. Other tools need compatible reporters/adapters; being listed is not support. No zero/not-run pass. Read `references/mutation-evidence.md` for setup, source-scope policy and limits.
## Setup
1. Copy this skill's `scripts/` into `<repo>/scripts/gauntlet/`; keep helpers together. Ignore `.crap/` (gate environments/exports) and `.gauntlet/` (run state). **Commit the scripts and ignores** before `start`: fresh worktrees see only committed files. The scorer `crap.py` is vendored from [crap-check](https://github.com/n0an/crap-check), pinned in its header; no second install is needed. With crap-check installed, cleaner can also use its repair loop.
2. Require Python 3.8+ for mutation evidence. Swift needs its toolchain and Muter v16 (`brew install muter-mutation-testing/formulae/muter`, verify version). Optional: `brew install swiftlint` for structure rules.
3. Gate a module: `scripts/gauntlet/crap-gate.sh <ModuleName>` runs lizard, `swift test --enable-code-coverage`, `llvm-cov export` (never `--summary-only`, it strips per-function data) and the scorer. Env: `GAUNTLET_MODULES_DIR` (default `Modules`), `GAUNTLET_CRAP_THRESHOLD` (6), `GAUNTLET_COV_FLOOR` (70).
**On an app project, SPM mode cannot run - use `--xcresult`.** `swift test` builds nothing for an app target, and nothing for a package whose dependencies are iOS-only, so there is no profdata for llvm-cov. Such a project tests through the simulator, and its coverage lives in an `.xcresult`:
```bash
xcodebuild test -workspace App.xcworkspace -scheme App \
-destination 'platform=iOS Simulator,name=iPhone 17 Pro' \
-enableCodeCoverage YES -resultBundlePath .crap/cov.xcresult
scripts/gauntlet/crap-gate.sh --xcresult .crap/cov.xcresult Packages/Networking/Sources
```
Source directories also filter `xccov` paths to avoid per-file calls across the whole app; override with `GAUNTLET_XCCOV_INCLUDE`. In `.xcresult` mode `GAUNTLET_COV_FLOOR` is ignored: per-function CRAP is the gate.
4. Mutation: follow the adapter setup in `references/mutation-evidence.md`. Commit config, source and tests; capture with `mutation-evidence.py run-muter` (Swift) or `run-report` (standard JSON). Both bind to printed `MUTATION_BASE` and HEAD. Then hand off with regular tests; never relabel captures after commits.
Exit codes: 0 passed, 1 something over threshold, 2 could not measure (missing sources, red tests, no coverage export). A function flagged `not in the coverage report` means the test run never loaded that file - a build problem, not a testing gap.
## Practical notes
- Keep stories small (one story, 3-7 scenarios). Cost of change is ~zero; iterate instead of plan-maxing.
- Specs are **throwaway scaffolding**, not documentation: the feature file lives for one story cycle; the code is the only source of truth. Don't build a spec↔code sync habit - it never pays off.
- Prefer running the loop in fast, host-testable modules (SPM packages, workspace packages): mutation testing rebuilds per mutant, so test-loop speed decides whether the hardener is affordable. Each worktree also pays its own cold build and `.crap/venv`; that is the price of isolation, so `finish` runs as soon as the branch is reviewed.
- A stage that went wrong is restarted with `stage <slug> <stage> --retry` (the failed attempt stays under `refs/gauntlet/<slug>/<stage>/attempt-N`); a lost orchestrator recovers with `status`.
- In Claude Code with this repo installed as a plugin, the whole pipeline is one command: `/gauntlet <story text>` (add `--qa` for phase 2). Elsewhere, any harness that can run staged sessions drives the same five `gauntlet.sh` commands by hand.