Installs into .claude/skills of the current project.
Are you the author of Cleanup Loops?
Add the live security badge to your README. It updates with every re-scan.
[](https://www.skillsdirectory.com/skills/brennontwilliams-cleanup-loops)
---
name: cleanup-loops
description: Use when asked to clean up stuck loops, kill dead loop processes, or troubleshoot loop state.
disable-model-invocation: true
argument-hint: "[--dry-run] [--threshold N] [--interrupted-age H]"
model: sonnet
allowed-tools:
- Bash(ll-loop:*, kill:*, rm:*, python3:*)
- AskUserQuestion
arguments:
- name: dry_run
description: Preview discovered stuck/stale loops without cleaning anything (default false)
required: false
- name: threshold
description: Minutes before a "running" loop's updated_at is considered stale (default 15)
required: false
- name: interrupted_age
description: Hours before an "interrupted" loop is considered abandoned and offered for cleanup, regardless of PID artifacts (default 24)
required: false
metadata:
short-description: Use when asked to clean up stuck loops, kill dead loop processes, or troubleshoo
---
# Cleanup Loops
Discover stuck or stale `ll-loop` processes, diagnose root causes from their state and events
files, and clean them up after user confirmation.
---
## Step 1: Enumerate All Loops with State Files
```bash
ll-loop list --all-runs --json
```
This returns a JSON array of all loops that have a `.state.json` file in `.loops/.running/`,
regardless of whether they are actually running. Each element contains:
| Field | Type | Description |
|---|---|---|
| `loop_name` | string | Unique loop identifier |
| `instance_id` | string | Per-instance timestamp stem (e.g. `fix-types-20260503T122306`); absent from list output — use `ll-loop status <loop_name> --json` to resolve |
| `status` | string | `"running"`, `"interrupted"`, `"failed"`, `"timed_out"`, `"awaiting_continuation"`, `"completed"` |
| `current_state` | string | Last active FSM state name |
| `updated_at` | string | ISO 8601 UTC timestamp of last state save |
| `accumulated_ms` | integer | Total elapsed milliseconds |
| `iteration` | integer | Last iteration count |
| `last_result` | object\|null | Last evaluation verdict and details |
If the command fails or returns an empty array (`[]`), report:
```
No loop state files found. Nothing to clean.
```
and stop.
---
## Step 2: Gather Detailed Status for Each Loop
For each loop returned in Step 1, run:
```bash
ll-loop status <loop_name> --json 2>/dev/null
```
> **Note**: `ll-loop status` performs first-pass reconciliation (ENH-1669). If a state file claims `running` but its PID is provably dead, it is automatically rewritten to `interrupted` with a `reconciled_at` timestamp. When no PID is resolvable at all, a 6h `updated_at`-age fallback catches permanently PID-less orphans the same way (BUG-3317) — this skill's own 15-minute staleness check below still does non-redundant work inside that window, since it is a more aggressive, user-confirmed heuristic rather than an automatic one. This means loops that were orphaned foreground crashes will already show `interrupted` by the time you reach Step 3, reducing the number of manual cleanup actions required.
This produces the same fields as Step 1 plus two additional fields:
| Field | Type | Description |
|---|---|---|
| `pid` | integer\|null | PID from `.loops/.running/<instance_id>.pid` **or** `.lock` file; null if neither exists |
| `pid_source` | string\|null | `"pid_file"` if PID came from `.pid` file; `"lock_file"` if from `.lock` file; null if no PID |
Also check whether the PID is alive (only when `pid` is non-null):
```bash
kill -0 <pid> 2>/dev/null && echo "alive" || echo "dead"
```
---
## Step 3: Classify Each Loop
Use the current UTC time and the loop's `updated_at` to compute staleness. There are two
thresholds:
- **Running-stale threshold**: `${threshold:-15}` minutes — applied to `status == "running"`
loops to detect dead/abandoned foreground runs.
- **Interrupted-aged threshold**: `${interrupted_age:-24}` hours — applied to
`status == "interrupted"` loops with no orphaned PID/lock file, to surface accumulated
clean-Ctrl-C runs that the user is unlikely to resume.
To compute age in minutes from an ISO 8601 timestamp:
```bash
python3 -c "
from datetime import datetime, timezone
updated = datetime.fromisoformat('$UPDATED_AT'.replace('Z', '+00:00'))
now = datetime.now(timezone.utc)
print(f'{(now - updated).total_seconds() / 60:.1f}')
"
```
Classify each loop into one of the following categories:
### NEEDS CLEANUP — stuck-running
**Condition**: `status == "running"` AND either:
- `pid` is non-null AND the process is dead (`kill -0` returned non-zero), OR
- `updated_at` is older than the threshold
**Action**: Call `ll-loop stop <loop_name>`.
`ll-loop stop` handles all cleanup: sends SIGTERM (then SIGKILL after 10s) if the PID is
still alive, updates `status` to `interrupted`, and deletes the `.pid` file.
### NEEDS CLEANUP — stale-interrupted
**Condition**: `status == "interrupted"` AND `pid` is non-null (orphaned artifact present)
**Note**: Interrupted loops are resumable — offer `ll-loop resume` before cleanup:
```bash
ll-loop resume <loop_name> # preferred: pick up where the loop left off
```
If the user wants to discard the run and only needs the orphaned artifact removed, then
remove the stale artifact based on `pid_source`:
- `pid_source == "pid_file"` → remove the `.pid` file:
```bash
rm -f ".loops/.running/<instance_id_or_loop_name>.pid"
```
- `pid_source == "lock_file"` → remove the `.lock` file:
```bash
rm -f ".loops/.running/<instance_id_or_loop_name>.lock"
```
The `.state.json` is preserved for diagnostics.
### NEEDS CLEANUP — stale-interrupted-aged
**Condition**: `status == "interrupted"` AND `updated_at` older than `${interrupted_age:-24}` hours AND
no orphaned PID/lock file (those fall under `stale-interrupted` above).
**Why this exists**: Clean Ctrl-C exits leave a state file but no orphaned PID, so they don't
qualify as `stale-interrupted`. Without aging-based detection, these accumulate indefinitely —
e.g. a long-running loop like `autodev` that the user routinely interrupts and never resumes
piles up dozens of `.state.json` files in `.loops/.running/`. After a day, the run is
overwhelmingly unlikely to be resumed.
**Note**: Still resumable in principle — offer `ll-loop resume` before archiving:
```bash
ll-loop resume <loop_name> # preferred if the work is still relevant
```
If the user wants to discard the run, archive the state and events files to `.loops/.history/`
and remove them from `.running/`. The cleanest path is to use the same archival mechanism the
startup sweep uses (`StatePersistence.clear_all()`); if that's awkward to invoke from a skill,
fall back to a direct copy-then-delete using the `<run_id>-<loop_name>` convention:
```bash
run_id=$(python3 -c "
import json
d = json.load(open('.loops/.running/<instance_id>.state.json'))
print(d['started_at'].replace(':','').replace('.','').replace('+','')[:17]
)")
mkdir -p ".loops/.history/${run_id}-<loop_name>"
mv .loops/.running/<instance_id>.state.json ".loops/.history/${run_id}-<loop_name>/state.json"
mv .loops/.running/<instance_id>.events.jsonl ".loops/.history/${run_id}-<loop_name>/events.jsonl" 2>/dev/null || true
```
The `.state.json` and `.events.jsonl` end up in `.history/` for diagnostics; `.running/` is freed.
### NEEDS ATTENTION — abandoned-handoff
**Condition**: `status == "awaiting_continuation"` AND `updated_at` older than the threshold
**Action**: Surface to user in the summary. Do NOT auto-clean; the user may want to resume.
### INFORMATIONAL — terminal
**Condition**: `status` is `"failed"` or `"timed_out"`
**Action**: Report in the summary. These are already in terminal states; their state files
are diagnostic artifacts. Do NOT auto-clean.
### HEALTHY — skip
**Condition**: `status == "running"` AND `pid` alive AND `updated_at` is fresh
**Condition**: `status == "awaiting_continuation"` AND `updated_at` is fresh
**Condition**: `status == "completed"`
**Action**: Skip — leave these loops alone.
---
## Step 4: Display Summary
Print a summary of all discovered loops grouped by category:
```
Loop State Summary
==================
NEEDS CLEANUP (N):
[1] <loop_name> — stuck-running — status: running — PID: 12345 (dead) — last updated: 47m ago
[2] <loop_name> — stale-interrupted — status: interrupted — stale PID file — last updated: 2h ago
[3] <loop_name> — stale-interrupted-aged — status: interrupted — no PID artifact — last updated: 3d ago
NEEDS ATTENTION (M):
[3] <loop_name> — abandoned-handoff — status: awaiting_continuation — last updated: 3h ago
INFORMATIONAL (K):
[4] <loop_name> — terminal — status: failed — last updated: 1h ago
HEALTHY (J):
<loop_name> — running — PID 9876 (alive) — last updated: 2m ago
```
If there are no loops in any actionable category (NEEDS CLEANUP or NEEDS ATTENTION), report:
```
No stuck or stale loops found.
<HEALTHY section or "No loops with state files.">
```
and stop.
If `--dry-run` flag is set, stop here — do not proceed to cleanup.
---
## Step 5: Confirm Cleanup
Use `AskUserQuestion` to ask which loops to clean up. Include only loops numbered under
NEEDS CLEANUP in the prompt (NEEDS ATTENTION loops are informational — let the user
address them manually):
```
Clean up N loop(s) marked NEEDS CLEANUP? [Y/n/select]
Y — clean up all N loops
n — cancel, no changes made
select — choose which (enter comma-separated numbers from the list above, e.g. 1,2)
```
Handle the response:
- `Y` or empty/enter → clean up all NEEDS CLEANUP loops
- `n` → report "No changes made." and stop
- `select` or comma-separated numbers (e.g. `1,3`) → clean only the listed indices
---
## Step 6: Execute Cleanup
For each loop confirmed for cleanup:
### stuck-running loops
```bash
ll-loop stop <loop_name>
```
If `ll-loop stop` exits with a non-zero code (e.g. the process already died between steps),
report:
```
[WARN] ll-loop stop exited with error for <loop_name>; loop may have already terminated.
```
### stale-interrupted-aged loops
First offer to resume the loop — even aged interrupts are still resumable in principle:
```bash
ll-loop resume <loop_name>
```
If the user wants to archive and free the slot, move the state + events to `.history/`:
```bash
state_file=$(ls .loops/.running/<loop_name>-*.state.json 2>/dev/null | sort | tail -1)
[ -n "$state_file" ] || { echo "no state file for <loop_name>"; continue; }
stem=$(basename "$state_file" .state.json)
run_id=$(python3 -c "
import json
d = json.load(open('$state_file'))
print(d['started_at'].replace(':','').replace('.','').replace('+','')[:17]
)")
archive_dir=".loops/.history/${run_id}-<loop_name>"
mkdir -p "$archive_dir"
mv "$state_file" "$archive_dir/state.json"
mv ".loops/.running/${stem}.events.jsonl" "$archive_dir/events.jsonl" 2>/dev/null || true
echo "Archived $stem to $archive_dir"
```
### stale-interrupted loops
First offer to resume the loop — interrupted loops are resumable:
```bash
ll-loop resume <loop_name> # preferred: pick up where the loop left off
```
If the user wants to discard the run and only needs the orphaned artifact removed,
branch on `pid_source` from Step 2:
- `pid_source == "pid_file"`:
```bash
rm -f ".loops/.running/<instance_id_or_loop_name>.pid"
echo "Removed stale .pid file for <loop_name>"
```
- `pid_source == "lock_file"`: check whether the lock-holder process is still alive:
- **PID alive** (`kill -0 <pid>` exits 0): the process is an orphaned lock holder
blocking scope acquisition. Use `ll-loop stop` to kill it and remove the lock:
```bash
ll-loop stop <loop_name>
```
- **PID dead** (`kill -0 <pid>` exits non-zero): the process is gone but left a
stale file. Remove it directly:
```bash
rm -f ".loops/.running/<instance_id_or_loop_name>.lock"
echo "Removed stale .lock file for <loop_name>"
```
---
## Step 7: Inspect Root Cause
For every cleaned loop (both stuck-running and stale-interrupted), tail the events file to
surface what happened just before the failure:
```bash
f=$(ls .loops/.running/<loop_name>-*.events.jsonl 2>/dev/null | sort | tail -1)
[ -n "$f" ] && tail -20 "$f"
```
If the events file is missing, note: `(no events file found — loop may not have started)`
Parse and format the last events as a brief timeline. Look for patterns like:
- Last event type (e.g. `state_enter`, `action_complete`, `evaluate`)
- Any `exit_code` != 0 in `action_complete` events
- Any `verdict == "fail"` in `evaluate` events
- Any `terminated_by` values in `loop_complete` events
---
## Step 8: Final Report
For each cleaned loop, display a block:
```
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
Loop: <loop_name> [CLEANED]
Was stuck in: <current_state>
Status: <status> → interrupted
Last updated: <updated_at> (<N> minutes ago)
Iteration: <iteration>
PID: <pid> (dead) / none
Last events:
<ts> state_enter state=<state>
<ts> action_complete exit_code=<N> duration=<N>ms
...
Root cause: <brief interpretation — e.g. "Action exited non-zero in <state> state with
no further events, suggesting the subprocess crashed or was killed externally.">
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
```
After all blocks, print a summary line:
```
Cleaned <N> loop(s). <M> loop(s) need attention (see NEEDS ATTENTION above).
```
---
## Usage Examples
```bash
# Discover and clean all stuck/stale loops (with confirmation)
/ll:cleanup-loops
# Preview what would be cleaned without making changes
/ll:cleanup-loops --dry-run
# Use a custom staleness threshold (30 minutes instead of default 15)
/ll:cleanup-loops --threshold 30
# Dry run with custom threshold
/ll:cleanup-loops --dry-run --threshold 60
# Prune aged interrupted loops more aggressively (after 6h instead of default 24h)
/ll:cleanup-loops --interrupted-age 6
# Conservative: only flag interrupts that are at least a week old
/ll:cleanup-loops --interrupted-age 168
```