Drive dirtywork — delegate a coding task to a local model that works in an isolated git worktree, then review the result, resume it with feedback, and record a verdict. Use when asked to run dirtywork, hand implementation to a local/worker model, or review a dirtywork run.
Installs into .claude/skills of the current project.
Are you the author of Contract?
Add the live security badge to your README. It updates with every re-scan.
[](https://www.skillsdirectory.com/skills/jimboschneider-contract)
---
name: dirtywork
description: Drive dirtywork — delegate a coding task to a local model that works in an isolated git worktree, then review the result, resume it with feedback, and record a verdict. Use when asked to run dirtywork, hand implementation to a local/worker model, or review a dirtywork run.
---
# Driving dirtywork (v{{VERSION}})
*If you are the dirtywork **worker** reading this inside the sandbox: this file
is for the orchestrator that launched you. Ignore it and follow your task.*
You are the orchestrator. dirtywork runs a **worker** — a local model — in an
agentic tool-use loop inside an isolated git worktree and hands you the result.
You pick the task, write the brief, review what comes back, and decide. The
worker never merges anything; you do.
**Reference:** `dirtywork contract` prints every flag, the stdout JSON schema,
exit codes and transcript events for the installed version. Run it before your
first `run` in a session. Do not guess flags.
## Before the first run
- `dirtywork --version` — installed and on PATH (`pipx install dirtywork` if not).
- The endpoint you will pass is serving and the model you will pass as `--model`
is loaded (LM Studio `curl -s http://localhost:1234/v1/models`; Ollama
`ollama ps`) or available (another OpenAI-compatible server:
`curl -s <base-url>/models`). If the request does not say which provider,
find out; if more than one is running, ask — never pick one silently, and
never assume LM Studio. `--provider anthropic` needs `ANTHROPIC_API_KEY`;
an OpenAI-compatible server with a key needs `OPENAI_API_KEY` (sent as a
Bearer token when set).
- `docker info` — docker mode (the default, and the contained one) needs a
running Docker; the worker image is pulled on the first run if it is
absent (exit 2 with "Build or pull the worker image" means that failed —
see the contract's `--image` entry). `--sandbox none` runs the worker on
your host: read the contract's **Security** entry first.
## The loop
### 1. Write the brief
One run = one task. The worker is a small model with a small context window: it
executes, it does not design. Put in the brief:
- the exact files to touch, by path — and any it must not touch;
- the exact names, strings and signatures, quoted;
- the test(s) that must pass and the command that runs them (the worker runs
them itself);
- decisions already made ("use X, not Y") so it does not re-decide them.
Keep it under ~40 lines. If you cannot name the files, you are not ready to
delegate — explore first, then brief. The target repo's `CLAUDE.md` /
`AGENTS.md` (at the base commit) is injected into the worker's prompt
automatically: put worker-facing conventions there, not in every brief.
### 2. Run
dirtywork run --repo <path> "<brief>" \
--model <model> \
[--provider openai|anthropic|ollama] [--base-url <url>] \
--verify "<test command>" --verify-rounds 2 \
--max-turns 60 --timeout 1800
- `--provider`/`--base-url`: omit both for LM Studio on `localhost:1234` (the
default). Use `--provider ollama` for Ollama; for another OpenAI-compatible
server, pass `--base-url <url>`. `resume` inherits `--provider`, but it does
not restore a custom `--base-url` — repeat that custom URL on every resume.
- `--verify` runs your gate in the sandbox after the worker finishes and feeds
failures back for up to `--verify-rounds` further attempts.
- Progress is on **stderr**: transcript path, worktree path, `error:` lines.
`tail -f` the transcript to watch a run.
- **Past preflight, stdout is exactly one JSON object.** Parse it; do not
grep it. On exit 2 stdout is empty and the reason is the `error:` line on
stderr. Fields you will use: `status`, `worktree`, `branch`, `base_commit`,
`transcript`, `run_dir` (its last path component is the slug),
`resumed_from`, `turns`, `files_changed`, `final_message`, `stuck_on`,
`last_tool_result`, `last_assistant_text`.
- Exit codes: `0` = `completed`. `1` = any other status (`max_turns`,
`timeout`, `stalled`, `stuck`, `verify_failed`, `context_exhausted`,
`model_error`, `budget_exceeded`, `unchanged`, …) — the worktree and branch
are kept for review and salvage. `2` = preflight or environment error;
nothing was created.
- In the **default** docker mode the container has no network and no host
directories mounted (only the run volume and a read-only view of the repo's
git objects), so the worker **cannot install dependencies**; it has what the image
ships — git, bash, python3, node/npm, .NET, ripgrep, jq, curl, shellcheck.
`--allow-network` gives the container bridge networking so installs work and
the offline guarantee is gone — use it deliberately, and for anything
permanent build a derived image instead (see the contract). Run any richer
gate yourself, on the host, against the exported worktree.
### 3. Review — always, before anything merges
- `git -C <worktree> status` then `git -C <worktree> diff`: read the diff
first, the transcript second. Compare `files_changed` with the brief —
a touched file you did not name is a finding.
- Run the repo's own gate on the host against the worktree.
- Judge the work, not the status: `completed` with a wrong change is a reject;
`max_turns` with a correct change is salvageable.
### 4. Resume with feedback, or re-brief
If the work is close: `dirtywork resume <slug> --feedback "<exactly what to
change>"` — same worktree, same branch; the worker is told to inspect its
earlier work and apply your feedback. A resume is a **new run** with its own
slug and `run_dir`; its JSON's `resumed_from` names the run it continued while
`branch` and `worktree` still carry the original slug. From here on use the
newest slug for `runs verdict` and `runs clean` — cleaning an earlier slug in
the chain leaves the worktree and branch in place. Two resumes without
convergence means the brief was wrong: reject, rewrite the brief, run again
from a clean base.
### 5. Verdict and cleanup
- `dirtywork runs verdict <slug> accept|reject --note "<why>"` records your
decision in the run's `run.json`.
- Accept: commit or PR the `dirtywork/<slug>` branch as you would any
contributor's work.
- Reject: record it, then discard the work but not the record —
`dirtywork runs clean <slug> --force --keep-transcript`. `--force` because a
rejected worktree has uncommitted changes and `clean` refuses to delete
those otherwise; `--keep-transcript` keeps `run.json` and the transcript and
removes the container, volume, worktree and branch. Plain `runs clean <slug>`
deletes the receipts too.
- `dirtywork runs list` and `dirtywork runs show <slug>` inspect earlier runs.
## Rules of thumb
- Never merge unreviewed worker output. The worker ran `bash` in a sandbox;
a shell is a shell — see the contract's **Security** entry.
- Parallelism is processes: run independent briefs concurrently (LM Studio
serves several requests per loaded model). Never two tasks in one brief.
- On `stuck`, read `stuck_on` (the repeated failing command); on `stalled` or
`max_turns`, read `last_tool_result` and `last_assistant_text`. Either way
the fix is usually in the brief, not the model.
- Keep receipts: `run.json`, the transcript and the verdict are the record of
what the worker did and what you decided. `runs clean --keep-transcript`
preserves them; plain `clean` does not.