Skip to content
Back to skills

Cleanup Loops

ASecurity

Use when asked to clean up stuck loops, kill dead loop processes, or troubleshoot loop state.

  • 6 stars
  • 0 votes
  • 0 copies
  • 0 views
  • Added October 6, 2026
ai-agentspythongobash

Works with

  • terminal

Security analysis

A100/100

Pro scans all 2 files and shows the line behind each finding

Scanned October 6, 2026

npx -y skills add BrennonTWilliams/little-loops --skill cleanup-loops --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Cleanup Loops?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Cleanup Loops
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/brennontwilliams-cleanup-loops/badge)](https://www.skillsdirectory.com/skills/brennontwilliams-cleanup-loops)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: cleanup-loops
description: Use when asked to clean up stuck loops, kill dead loop processes, or troubleshoot loop state.
disable-model-invocation: true
argument-hint: "[--dry-run] [--threshold N] [--interrupted-age H]"
model: sonnet
allowed-tools:
  - Bash(ll-loop:*, kill:*, rm:*, python3:*)
  - AskUserQuestion
arguments:
  - name: dry_run
    description: Preview discovered stuck/stale loops without cleaning anything (default false)
    required: false
  - name: threshold
    description: Minutes before a "running" loop's updated_at is considered stale (default 15)
    required: false
  - name: interrupted_age
    description: Hours before an "interrupted" loop is considered abandoned and offered for cleanup, regardless of PID artifacts (default 24)
    required: false
metadata:
  short-description: Use when asked to clean up stuck loops, kill dead loop processes, or troubleshoo
---

# Cleanup Loops

Discover stuck or stale `ll-loop` processes, diagnose root causes from their state and events
files, and clean them up after user confirmation.

---

## Step 1: Enumerate All Loops with State Files

```bash
ll-loop list --all-runs --json
```

This returns a JSON array of all loops that have a `.state.json` file in `.loops/.running/`,
regardless of whether they are actually running. Each element contains:

| Field | Type | Description |
|---|---|---|
| `loop_name` | string | Unique loop identifier |
| `instance_id` | string | Per-instance timestamp stem (e.g. `fix-types-20260503T122306`); absent from list output — use `ll-loop status <loop_name> --json` to resolve |
| `status` | string | `"running"`, `"interrupted"`, `"failed"`, `"timed_out"`, `"awaiting_continuation"`, `"completed"` |
| `current_state` | string | Last active FSM state name |
| `updated_at` | string | ISO 8601 UTC timestamp of last state save |
| `accumulated_ms` | integer | Total elapsed milliseconds |
| `iteration` | integer | Last iteration count |
| `last_result` | object\|null | Last evaluation verdict and details |

If the command fails or returns an empty array (`[]`), report:

```
No loop state files found. Nothing to clean.
```

and stop.

---

## Step 2: Gather Detailed Status for Each Loop

For each loop returned in Step 1, run:

```bash
ll-loop status <loop_name> --json 2>/dev/null
```

> **Note**: `ll-loop status` performs first-pass reconciliation (ENH-1669). If a state file claims `running` but its PID is provably dead, it is automatically rewritten to `interrupted` with a `reconciled_at` timestamp. When no PID is resolvable at all, a 6h `updated_at`-age fallback catches permanently PID-less orphans the same way (BUG-3317) — this skill's own 15-minute staleness check below still does non-redundant work inside that window, since it is a more aggressive, user-confirmed heuristic rather than an automatic one. This means loops that were orphaned foreground crashes will already show `interrupted` by the time you reach Step 3, reducing the number of manual cleanup actions required.

This produces the same fields as Step 1 plus two additional fields:

| Field | Type | Description |
|---|---|---|
| `pid` | integer\|null | PID from `.loops/.running/<instance_id>.pid` **or** `.lock` file; null if neither exists |
| `pid_source` | string\|null | `"pid_file"` if PID came from `.pid` file; `"lock_file"` if from `.lock` file; null if no PID |

Also check whether the PID is alive (only when `pid` is non-null):

```bash
kill -0 <pid> 2>/dev/null && echo "alive" || echo "dead"
```

---

## Step 3: Classify Each Loop

Use the current UTC time and the loop's `updated_at` to compute staleness. There are two
thresholds:

- **Running-stale threshold**: `${threshold:-15}` minutes — applied to `status == "running"`
  loops to detect dead/abandoned foreground runs.
- **Interrupted-aged threshold**: `${interrupted_age:-24}` hours — applied to
  `status == "interrupted"` loops with no orphaned PID/lock file, to surface accumulated
  clean-Ctrl-C runs that the user is unlikely to resume.

To compute age in minutes from an ISO 8601 timestamp:

```bash
python3 -c "
from datetime import datetime, timezone
updated = datetime.fromisoformat('$UPDATED_AT'.replace('Z', '+00:00'))
now = datetime.now(timezone.utc)
print(f'{(now - updated).total_seconds() / 60:.1f}')
"
```

Classify each loop into one of the following categories:

### NEEDS CLEANUP — stuck-running

**Condition**: `status == "running"` AND either:
- `pid` is non-null AND the process is dead (`kill -0` returned non-zero), OR
- `updated_at` is older than the threshold

**Action**: Call `ll-loop stop <loop_name>`.
`ll-loop stop` handles all cleanup: sends SIGTERM (then SIGKILL after 10s) if the PID is
still alive, updates `status` to `interrupted`, and deletes the `.pid` file.

### NEEDS CLEANUP — stale-interrupted

**Condition**: `status == "interrupted"` AND `pid` is non-null (orphaned artifact present)

**Note**: Interrupted loops are resumable — offer `ll-loop resume` before cleanup:
```bash
ll-loop resume <loop_name>   # preferred: pick up where the loop left off
```

If the user wants to discard the run and only needs the orphaned artifact removed, then
remove the stale artifact based on `pid_source`:
- `pid_source == "pid_file"` → remove the `.pid` file:
  ```bash
  rm -f ".loops/.running/<instance_id_or_loop_name>.pid"
  ```
- `pid_source == "lock_file"` → remove the `.lock` file:
  ```bash
  rm -f ".loops/.running/<instance_id_or_loop_name>.lock"
  ```

The `.state.json` is preserved for diagnostics.

### NEEDS CLEANUP — stale-interrupted-aged

**Condition**: `status == "interrupted"` AND `updated_at` older than `${interrupted_age:-24}` hours AND
no orphaned PID/lock file (those fall under `stale-interrupted` above).

**Why this exists**: Clean Ctrl-C exits leave a state file but no orphaned PID, so they don't
qualify as `stale-interrupted`. Without aging-based detection, these accumulate indefinitely —
e.g. a long-running loop like `autodev` that the user routinely interrupts and never resumes
piles up dozens of `.state.json` files in `.loops/.running/`. After a day, the run is
overwhelmingly unlikely to be resumed.

**Note**: Still resumable in principle — offer `ll-loop resume` before archiving:
```bash
ll-loop resume <loop_name>   # preferred if the work is still relevant
```

If the user wants to discard the run, archive the state and events files to `.loops/.history/`
and remove them from `.running/`. The cleanest path is to use the same archival mechanism the
startup sweep uses (`StatePersistence.clear_all()`); if that's awkward to invoke from a skill,
fall back to a direct copy-then-delete using the `<run_id>-<loop_name>` convention:

```bash
run_id=$(python3 -c "
import json
d = json.load(open('.loops/.running/<instance_id>.state.json'))
print(d['started_at'].replace(':','').replace('.','').replace('+','')[:17]
)")
mkdir -p ".loops/.history/${run_id}-<loop_name>"
mv .loops/.running/<instance_id>.state.json    ".loops/.history/${run_id}-<loop_name>/state.json"
mv .loops/.running/<instance_id>.events.jsonl  ".loops/.history/${run_id}-<loop_name>/events.jsonl" 2>/dev/null || true
```

The `.state.json` and `.events.jsonl` end up in `.history/` for diagnostics; `.running/` is freed.

### NEEDS ATTENTION — abandoned-handoff

**Condition**: `status == "awaiting_continuation"` AND `updated_at` older than the threshold

**Action**: Surface to user in the summary. Do NOT auto-clean; the user may want to resume.

### INFORMATIONAL — terminal

**Condition**: `status` is `"failed"` or `"timed_out"`

**Action**: Report in the summary. These are already in terminal states; their state files
are diagnostic artifacts. Do NOT auto-clean.

### HEALTHY — skip

**Condition**: `status == "running"` AND `pid` alive AND `updated_at` is fresh
**Condition**: `status == "awaiting_continuation"` AND `updated_at` is fresh
**Condition**: `status == "completed"`

**Action**: Skip — leave these loops alone.

---

## Step 4: Display Summary

Print a summary of all discovered loops grouped by category:

```
Loop State Summary
==================

NEEDS CLEANUP (N):
  [1] <loop_name> — stuck-running — status: running — PID: 12345 (dead) — last updated: 47m ago
  [2] <loop_name> — stale-interrupted — status: interrupted — stale PID file — last updated: 2h ago
  [3] <loop_name> — stale-interrupted-aged — status: interrupted — no PID artifact — last updated: 3d ago

NEEDS ATTENTION (M):
  [3] <loop_name> — abandoned-handoff — status: awaiting_continuation — last updated: 3h ago

INFORMATIONAL (K):
  [4] <loop_name> — terminal — status: failed — last updated: 1h ago

HEALTHY (J):
  <loop_name> — running — PID 9876 (alive) — last updated: 2m ago
```

If there are no loops in any actionable category (NEEDS CLEANUP or NEEDS ATTENTION), report:

```
No stuck or stale loops found.
<HEALTHY section or "No loops with state files.">
```

and stop.

If `--dry-run` flag is set, stop here — do not proceed to cleanup.

---

## Step 5: Confirm Cleanup

Use `AskUserQuestion` to ask which loops to clean up. Include only loops numbered under
NEEDS CLEANUP in the prompt (NEEDS ATTENTION loops are informational — let the user
address them manually):

```
Clean up N loop(s) marked NEEDS CLEANUP? [Y/n/select]

  Y      — clean up all N loops
  n      — cancel, no changes made
  select — choose which (enter comma-separated numbers from the list above, e.g. 1,2)
```

Handle the response:
- `Y` or empty/enter → clean up all NEEDS CLEANUP loops
- `n` → report "No changes made." and stop
- `select` or comma-separated numbers (e.g. `1,3`) → clean only the listed indices

---

## Step 6: Execute Cleanup

For each loop confirmed for cleanup:

### stuck-running loops

```bash
ll-loop stop <loop_name>
```

If `ll-loop stop` exits with a non-zero code (e.g. the process already died between steps),
report:
```
  [WARN] ll-loop stop exited with error for <loop_name>; loop may have already terminated.
```

### stale-interrupted-aged loops

First offer to resume the loop — even aged interrupts are still resumable in principle:
```bash
ll-loop resume <loop_name>
```

If the user wants to archive and free the slot, move the state + events to `.history/`:

```bash
state_file=$(ls .loops/.running/<loop_name>-*.state.json 2>/dev/null | sort | tail -1)
[ -n "$state_file" ] || { echo "no state file for <loop_name>"; continue; }
stem=$(basename "$state_file" .state.json)
run_id=$(python3 -c "
import json
d = json.load(open('$state_file'))
print(d['started_at'].replace(':','').replace('.','').replace('+','')[:17]
)")
archive_dir=".loops/.history/${run_id}-<loop_name>"
mkdir -p "$archive_dir"
mv "$state_file"                                    "$archive_dir/state.json"
mv ".loops/.running/${stem}.events.jsonl"          "$archive_dir/events.jsonl" 2>/dev/null || true
echo "Archived $stem to $archive_dir"
```

### stale-interrupted loops

First offer to resume the loop — interrupted loops are resumable:
```bash
ll-loop resume <loop_name>   # preferred: pick up where the loop left off
```

If the user wants to discard the run and only needs the orphaned artifact removed,
branch on `pid_source` from Step 2:

- `pid_source == "pid_file"`:
  ```bash
  rm -f ".loops/.running/<instance_id_or_loop_name>.pid"
  echo "Removed stale .pid file for <loop_name>"
  ```
- `pid_source == "lock_file"`: check whether the lock-holder process is still alive:
  - **PID alive** (`kill -0 <pid>` exits 0): the process is an orphaned lock holder
    blocking scope acquisition. Use `ll-loop stop` to kill it and remove the lock:
    ```bash
    ll-loop stop <loop_name>
    ```
  - **PID dead** (`kill -0 <pid>` exits non-zero): the process is gone but left a
    stale file. Remove it directly:
    ```bash
    rm -f ".loops/.running/<instance_id_or_loop_name>.lock"
    echo "Removed stale .lock file for <loop_name>"
    ```

---

## Step 7: Inspect Root Cause

For every cleaned loop (both stuck-running and stale-interrupted), tail the events file to
surface what happened just before the failure:

```bash
f=$(ls .loops/.running/<loop_name>-*.events.jsonl 2>/dev/null | sort | tail -1)
[ -n "$f" ] && tail -20 "$f"
```

If the events file is missing, note: `(no events file found — loop may not have started)`

Parse and format the last events as a brief timeline. Look for patterns like:
- Last event type (e.g. `state_enter`, `action_complete`, `evaluate`)
- Any `exit_code` != 0 in `action_complete` events
- Any `verdict == "fail"` in `evaluate` events
- Any `terminated_by` values in `loop_complete` events

---

## Step 8: Final Report

For each cleaned loop, display a block:

```
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
Loop: <loop_name>   [CLEANED]
  Was stuck in:  <current_state>
  Status:        <status> → interrupted
  Last updated:  <updated_at> (<N> minutes ago)
  Iteration:     <iteration>
  PID:           <pid> (dead) / none

  Last events:
    <ts>  state_enter  state=<state>
    <ts>  action_complete  exit_code=<N>  duration=<N>ms
    ...

  Root cause: <brief interpretation — e.g. "Action exited non-zero in <state> state with
              no further events, suggesting the subprocess crashed or was killed externally.">
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
```

After all blocks, print a summary line:

```
Cleaned <N> loop(s). <M> loop(s) need attention (see NEEDS ATTENTION above).
```

---

## Usage Examples

```bash
# Discover and clean all stuck/stale loops (with confirmation)
/ll:cleanup-loops

# Preview what would be cleaned without making changes
/ll:cleanup-loops --dry-run

# Use a custom staleness threshold (30 minutes instead of default 15)
/ll:cleanup-loops --threshold 30

# Dry run with custom threshold
/ll:cleanup-loops --dry-run --threshold 60

# Prune aged interrupted loops more aggressively (after 6h instead of default 24h)
/ll:cleanup-loops --interrupted-age 6

# Conservative: only flag interrupts that are at least a week old
/ll:cleanup-loops --interrupted-age 168
```

Files in this skill

  • SKILL.md13.9 KB
  • agents/openai.yaml147 B

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…