Skip to content
Back to skills

Executing Plans

ASecurity

Use IMMEDIATELY after writing-plans, OR whenever a plan exists in ~/skill-workspace/orchestrator.db. Owns ALL post-publish writes to the orchestrator database via the deterministic-flow scripts in scripts/ (run-step is the preferred path for shell work). Never bypass these scripts; never write to the DB ad-hoc. Triggers on every plan execution: code, SQL, schema, build, test, fix, implement, deploy, refactor, run, execute, continue, resume. When a test fails: never patch the test silently — f...

  • 26 stars
  • 0 votes
  • 0 copies
  • 0 views
  • Added October 10, 2026
ai-agentspythongoshellbashsqlnodedebuggingapidatabasedocumentation

Works with

  • terminal
  • cli
  • api

Security analysis

A100/100

Pro scans all 20 files and shows the line behind each finding

Scanned October 10, 2026

npx -y skills add yizhao95/prov_ledger --skill executing-plans --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Executing Plans?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Executing Plans
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/yizhao95-executing-plans/badge)](https://www.skillsdirectory.com/skills/yizhao95-executing-plans)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: executing-plans
description: "Use IMMEDIATELY after writing-plans, OR whenever a plan exists in ~/skill-workspace/orchestrator.db. Owns ALL post-publish writes to the orchestrator database via the deterministic-flow scripts in scripts/ (run-step is the preferred path for shell work). Never bypass these scripts; never write to the DB ad-hoc. Triggers on every plan execution: code, SQL, schema, build, test, fix, implement, deploy, refactor, run, execute, continue, resume. When a test fails: never patch the test silently — fail-step it, root-cause via systematic-debugging, then deviate via scripts/deviate.sh."
---

# Executing Plans

The deterministic write-flows (one from writing-plans, the rest here) are **the only way** to mutate `~/skill-workspace/orchestrator.db`. Compose a small JSON input file, run the matching script, done. Never invoke the orchestrator CLI directly for writes; never construct ad-hoc SQL.

## 🔌 Boundary with `writing-plans`

`writing-plans` does ONE thing: composes a `plan-input.{json,yaml}` file and runs `${CLAUDE_PLUGIN_ROOT}/skills/writing-plans/scripts/publish-plan.sh` to insert the new plan. After that, **every** subsequent operation on that plan (start, log, complete, fail, deviate, record deferred-load skills, finish) is owned by THIS skill via the scripts in `scripts/`.

If you find yourself wanting to "update the plan" — you ARE the update mechanism. You do not go back to writing-plans.

## 🛠️ The deterministic flows (cheat sheet)

| Script | When | Required input | Optional input |
|---|---|---|---|
| `scripts/run-step.sh` ⭐  | **Preferred for COMMAND/CODE/TEST** — bundles start+exec+complete with auto-captured log_context (migration 006 + Aug 2026) | `step_id`, `type`, `command` | `summary`, `allow_nonzero` |
| `scripts/start-step.sh`    | Manual start — use for THINKING/DOCUMENTATION/ANALYSIS or when you need to inspect intermediate output | `step_id` | `type`, `agent_input`, `log_context` |
| `scripts/complete-step.sh` | Manual complete after non-shell work, OR after a `start-step` you split from exec. **Not for `type=COMMAND`** (exit 5 without run-step's `--- exit_code=` footer) — shell work always goes through `run-step.sh` (S2/E6-2) | `step_id` | `summary`, `agent_output`, `log_context` |
| `scripts/fail-step.sh`     | Work failed and you can't recover in this step | `step_id` | `reason`, `log_context` |
| `scripts/append-log.sh`    | Captured raw shell/tool output worth keeping AFTER the fact | `step_id`, `text` | — |
| `scripts/deviate.sh`       | Realized the plan needs new sub-steps         | `parent_step_id`, `justification`, `sub_steps[]` | — |
| `scripts/record-skill.sh`  | Activated a NEW skill mid-flight              | `plan_id`, `name`, `source` | `step_id`, `reason` |
| `scripts/record-metric.sh` | Phase 7: a NUMERIC observation (a rollup's mean, an AUC) for the metric outcome channel — or pass `metrics_from_stdout: true` to `run-step.sh` and print `metric name=<x> value=<v> [unit=<u>]` | `name`, `value` (a number) | `unit`, `project` (defaults to the plan's), `plan_id`, `step_id` |
| `scripts/raise-budget.sh`  | The loop breaker refused a `deviate` at `revision_count`/`max_revisions` and the extra revisions are genuinely warranted | `plan_id`, `new_max` (an integer **above** the current ceiling), `reason` | — |
| `scripts/abandon-plan.sh` | A plan nobody started has to be put down — published by mistake, or left with no steps by a publish that failed half-way | `plan_id`, `reason` | — |
| `scripts/finish-plan.sh`   | **Usually auto** — manual only for back-fill of pre-2026-05-26 plans, or to force-finish a plan with PENDING steps | `plan_id` | — |
| `scripts/agent-review-close.sh` | **Only after a `needs_agent_review` handoff** — the review sub-agent's sole way to finalize a `NEEDS_REVIEW` plan | `plan_id`, `outcome` (`pass`\|`fail`) | `summary`, `log_context` |

### ⚡ `run-step.sh` — the preferred path for shell-based steps (Aug 2026)

Before Aug 2026, agents had to manually pass `log_context` on every `complete-step` to capture
stdout/stderr. A May 2026 audit found this was the #1 forgotten step — coverage had dropped
from 100% to 33%. `run-step.sh` eliminates the footgun the same way migration 006 eliminated
`finish-plan.sh`:

```json
{
  "step_id": "my-plan-A",
  "type":    "COMMAND",
  "command": "pytest tests/ -v"
}
```

```bash
bash scripts/run-step.sh /tmp/in.json
```

**What it does, atomically:**
1. Calls `start-step` with a `[run-step] kickoff <ts>` banner as initial `log_context`
2. Runs `bash -c "$command" 2>&1 | tee` — captures combined stdout+stderr (PIPESTATUS preserved)
3. Truncates output to ≤16 KiB (head 8 KiB + `--- TRUNCATED N BYTES ---` marker + tail 8 KiB)
4. Appends a `--- exit_code=N, runtime=Ns ---` footer
5. If exit 0 → calls `complete-step`; else → calls `fail-step`
6. Prints one JSON line last — `{"ok":true,"op":"run-step","exit_code":N,...}` — carrying every other key of the complete-step/fail-step tail line (so `needs_agent_review` reaches you here too), then exits with the wrapped command's true exit code (so the calling agent sees pass/fail naturally)
7. The auto-review trigger from migration 006 still fires — plans still auto-close

**Decision table for which entry point to use:**

| Situation | Use |
|---|---|
| Step IS a shell command / test / build / migration  | `run-step.sh` ⭐ |
| Step IS pure thinking / docs / analysis (no shell)  | `start-step.sh` + `complete-step.sh` |
| Multi-command shell work in one step                | `run-step.sh` with `"command": "cmd1 && cmd2"` or a heredoc |
| You need to inspect intermediate output before committing the step result | `start-step.sh`, then `complete-step.sh` with explicit `log_context` |


### 🤖 Auto-review-and-complete (migration 006 — May 2026)

Every newly-published plan now gets one extra terminal step with `is_review=1`
(step_id = `<plan>-REVIEW`). This step is flipped by a **deterministic Python
procedure** (`api.review_and_complete`), NOT by the LLM agent.

**The procedure auto-fires from `complete-step.sh` and `fail-step.sh`** any time
a non-review step transitions. Decision tree:

| Sibling state | Review step | Plan |
|---|---|---|
| Any PENDING / IN_PROGRESS | stays PENDING | stays IN_PROGRESS |
| All COMPLETED | → COMPLETED | → COMPLETED |
| All COMPLETED **and the plan mentions a registered project** | → **NEEDS_REVIEW** + a tracked child step `<plan>-REVIEW.1` is created | stays IN_PROGRESS (awaits agent) |
| All FAILED steps have **recovered** deviation sub-trees (rest terminal) | → COMPLETED | → COMPLETED |
| Any FAILED step is **unrecovered** (rest terminal) | → FAILED | → FAILED |

**Recovery semantics (Jun 2026):** a FAILED step is considered *recovered* when
every direct non-review child is itself recovered — recursively. So if step `D`
failed and you `deviate.sh`'d into `D.1` which COMPLETED, the plan no longer
fails because of `D`. A FAILED leaf step (no deviation children) is never
recovered and still poisons the plan. Recursion depth is bounded by the
`depth_level<=3` circuit breaker, so this can't pathologically deep-walk.

**Retrying a failed attempt (FL-138):** the rollup asks whether *every* sub-task
came through, so a **retry is a child of the attempt it retries** —
`deviate.sh` on `D.1` gives you `D.1.1`, never a second sibling `D.2`. A FAILED
`D.1` beside a COMPLETED `D.2` outvotes it for ever: `D` stays unrecovered, the
review never reopens and the plan cannot close. Nested, the recursion above
applies unchanged. Each retry costs one depth level, and the `depth_level<=3`
breaker allows two (`D.1.1`, `D.1.1.1`); past that, restructure rather than
reach for a sibling.

**Implications for the agent:**
- You almost never need to call `finish-plan.sh` anymore. The last `complete-step`
  on the last regular step closes the plan automatically.
- The auto-trigger ignores the review step itself (no recursion risk).
- For legacy plans without a review row (pre-006), `finish-plan.sh` falls back
  to the old unconditional COMPLETED behavior — your back-fill workflow is preserved.
- For a partial plan you want to force-finish (some steps still PENDING),
  `finish-plan.sh` also falls back to plain `complete_plan` as an escape hatch.

### 🤖 NEEDS_REVIEW handoff (registered-project plans)

When a plan **belongs to a registered project** — `Plans.project`, declared in
the plan-input or derived from the repo it was published from (FL-014, phase
3.5); the goal text no longer needs to mention the project, and a plan
published before the column existed is matched once from its goal and
labelled `legacy` — the deterministic procedure does **not** auto-close it. Instead it (1) flips the
review step to `NEEDS_REVIEW`, (2) bumps the plan revision, and (3) inserts a
**tracked child step** `<plan>-REVIEW.1` (type `SUB_AGENT`, status `PENDING`)
under the review step. The plan stays `IN_PROGRESS`, and the tail line of
`complete-step`/`fail-step` — and of `run-step.sh`, which passes it on — carries:

```json
{"ok":true,"op":"complete-step",...,"needs_agent_review":true,"project":"<name>","review_step_id":"<plan>-REVIEW","review_child_step_id":"<plan>-REVIEW.1"}
```

**When you see `needs_agent_review: true`, you MUST:**
1. `start-step` the child step `review_child_step_id` (`<plan>-REVIEW.1`).
2. Dispatch a sub-agent (Agent tool) that loads the **`update-project-state-graph`**
   skill, passing `plan_id`, `project`, and the child step id.
3. When the sub-agent finishes, finalize **the child step**:
   `complete-step` it on a clean review, or `fail-step` it on a gap. The plan
   then **auto-closes from the child outcome** — child COMPLETED → review +
   plan COMPLETED; child FAILED → review + plan FAILED. No extra close call.

**Do NOT** close the plan or the review step directly — drive the tracked child
step and let the deterministic procedure finalize. `agent-review-close.sh`
remains a documented **fallback** for plans without a child (e.g. legacy
in-flight plans) and keeps any child in sync. Plans mentioning no registered
project are unaffected and still auto-close as above.

Schema: [`update-input.schema.json`](update-input.schema.json) (oneOf, one per op).
Examples: [`update-input.example.json`](update-input.example.json) (one worked example per op).

### Invocation pattern (identical for every script)

```bash
# 1. compose input
cat > /tmp/in.json <<'EOF'
{"step_id": "deep-health-20260514153045-A", "type": "ANALYSIS"}
EOF

# 2. run script
bash ${CLAUDE_PLUGIN_ROOT}/skills/executing-plans/scripts/start-step.sh /tmp/in.json
```

The script:
- Validates required fields (clear error naming the offending field)
- Enforces state machine + circuit breakers via `orchestrator.api`
- Writes to `${ORCH_DB:-~/skill-workspace/orchestrator.db}`
- Prints the result as JSON to stdout, **followed by a single-line OK marker** as the final stdout line: `{"ok":true,"op":"<op>","step_id":"<id>"}`
- Exits non-zero on any validation/state failure (with `❌ apply_op: ...` to stderr)

### 🛡️ Safe invocation pattern (Bug 1 fix — read this!)

**⚠️ NEVER pipe through `| tail -1` without `2>&1` and an exit-code check.** Bash's pipefail in the script does NOT propagate to your outer pipeline; failures will be silently swallowed and your plan will desync from reality.

Use one of these instead:

```bash
# (a) Easiest — capture, then check the OK marker on the final line:
OUT=$(bash scripts/complete-step.sh /tmp/in.json) || { echo "FAILED: $OUT" >&2; exit 1; }
echo "$OUT" | tail -1 | python3 -c 'import json,sys; assert json.loads(sys.stdin.read())["ok"]'

# (b) If you really must pipe, redirect stderr too + use PIPESTATUS:
bash scripts/complete-step.sh /tmp/in.json 2>&1 | tail -1
[[ ${PIPESTATUS[0]} -eq 0 ]] || { echo "step write failed"; exit 1; }
```

## 📷 MANDATORY: capture raw output via `log_context` (Bug 2 fix)

For every COMMAND or CODE step, the raw shell/tool/test output **MUST** end up in `Steps.log_context` — either:

- **(preferred)** inline: pass `log_context: "<raw output>"` to `complete-step.sh` (or `start-step.sh` / `fail-step.sh`); OR
- **(fallback)** call `append-log.sh` separately before `complete-step.sh`.

A non-trivial step with `log_context = ""` is a process failure — the dashboard's log panel goes blank, and a plan whose steps carry no logs leaves nothing to debug from after the fact.

## 🚨 Mandatory rules

| Rule | Why |
|---|---|
| Use ONLY the scripts in `scripts/` for writes — never the CLI or ad-hoc SQL | The scripts validate every field and enforce the state machine; a hand-typed write skips both |
| `summary` is ONE human-curated sentence; `text` (in append-log) is raw machine output | Telemetry vs. curation are distinct fields |
| Once a step is COMPLETED, it is IMMUTABLE — `deviate.sh` on it is REJECTED (exit 4, `accepted:false`, breaker `soft`). To redo work, deviate on the **next non-terminal step** (or the plan's review step) and put the retry sub-step there | Audit trail |
| `revision_count` ≤ `max_revisions` (default 5) — circuit breaker. Past it, `deviate` is REFUSED; the only way to lift the ceiling is `scripts/raise-budget.sh` with a `reason` (it is recorded as a deviation on the plan) — never an `UPDATE` | Prevents thrash; and a raised budget is a decision, so it leaves a record |
| `depth_level` ≤ 3 — circuit breaker | Keeps plans auditable |
| For `SUB_AGENT` steps: capture both `agent_input` (in start-step) and `agent_output` (in complete-step) | The dashboard renders these in dedicated panels |
| For COMMAND/CODE steps: capture raw output via `log_context` (inline on complete-step, OR separate append-log call) — NOT optional in practice | The dashboard's log panel goes blank otherwise; post-hoc debugging fails |
| Never pipe through `\| tail -1` without `2>&1` + PIPESTATUS check — the wrapper's pipefail does NOT propagate to your outer shell | Silent failures will desync the plan from reality (the F-stuck-PENDING bug) |

## 👀 Reads stay direct

These don't mutate the plan, so you can call them however:

```bash
provledger plan <plan_id> [--step <step_id>] [--full]   # steps in tree order, their logs, the deviations
sqlite3 "${ORCH_DB:-$HOME/skill-workspace/orchestrator.db}" \
  "SELECT plan_id, status FROM Plans WHERE status='IN_PROGRESS';"
```

Every `provledger` read is listed in [`docs/cli.md`](../../docs/cli.md). The dashboard at http://localhost:8765 is the visual equivalent.

### 🧭 Close-time reasons: answer WHY, point at the words (DP phase 1)

When a registered-project plan closes, `reason-slots.sh` lists the data
points the plan changed. For **each node answer two questions: why not the
other way? who asked for it?** — not what the step did. Three answer shapes,
and the tier is decided by the shape, never by you:

| you send | tier | when |
|---|---|---|
| `{"node_key": "nk_…", "utterance_id": 42, "span": [0, 31]}` | **stated** | the user's recorded words say it — point at them (`reason-slots.sh` with `"draft": true` proposes spans) |
| `{"node_key": "nk_…", "interpretation": "…", "refs": [7]}` | **asserted** | your reading of it; `refs` are `reference` ids (an email, a ticket, a verbal exchange registered with `provledger note`) |
| `{"node_key": "nk_…", "unstated": true}` | **unstated** | you do not know — say so; never invent |

The old `{"node_key", "text"}` shape is refused (exit 2): free text cannot
become *stated*, whoever sends it. Ask ONCE for all slots; the user pressing
enter accepts the draft. `provledger-extensions.json` → `reasons.close_mode`
= `ask` (default) or `pending` (slots are recorded as unstated/unknown for a
later `provledger why --pending`).

- ✅ good: `{"node_key": "nk_1a2b", "utterance_id": 42, "span": [0, 41]}` where utterance 42 reads *"please keep fiscal weeks, finance reconciles weekly"* — the reason IS the user's sentence.
- ❌ bad: `{"node_key": "nk_1a2b", "interpretation": "changed load_orders to group by fiscal week"}` — that describes the change, not why it beat calendar weeks or who wanted it.

### 📰 Headline responses and `because` (DP phase 2)

| Script | When | Input |
|---|---|---|
| `scripts/headline-respond.sh` | a blocking finding in the plan headline needs an answer (or you proceed past it on purpose) | `{plan_id, finding_id, action: revise\|proceed, rationale?, cites?: [reason_id], by?: agent\|human}` — the finding's record and every cited id become `influence(via=headline_response)`; a second answer to the same finding exits 2 |

Every `reason-fill` answer shape may carry `"because": [reason_id, …]` — the
records this change leaned on (a constraint you kept, a rejected path you
avoided). Each id becomes `influence(via=reason_because)`, and the result
reports `adopted`. Cite what you actually used; `provledger why <qn>` prints the
ids.

**Shown is not read.** A record listed in a headline, a checklist or a hook's
injection is `read_hit`; only an id you cite is `influence`. Neither is derived
from the other, and no count anywhere claims that anyone *read* anything.

## 📚 Reference index

- [`reference/op-catalog.md`](reference/op-catalog.md) — full per-op reference (state transitions + common mistakes)
- [`reference/when-to-use.md`](reference/when-to-use.md) — the common wrong moves, and the right script for each
- [`update-input.schema.json`](update-input.schema.json) — formal schema (oneOf per op)
- [`update-input.example.json`](update-input.example.json) — copy-paste-edit examples

## 🧪 Tests

```bash
"${PROVLEDGER_VENV:-$HOME/skill-workspace/.venv}/bin/python" -m pytest \
  "${CLAUDE_PLUGIN_ROOT}/skills/executing-plans/tests" -q
```

Per-op tests (happy path, validation, state machine, circuit breakers), a full lifecycle smoke through the documented scripts, and `test_update_input_schema.py`, which holds the schema and examples to the code.

## 🔗 See also

- [`../../README.md`](../../README.md) — what provLedger does; [`../../INSTALL.md`](../../INSTALL.md) — installation, and the dashboard (§6)
- [`../writing-plans/SKILL.md`](../writing-plans/SKILL.md) — the other half of the contract; owns the initial publish flow
- **Dashboard** — the bundled `orchestrator-webapp/` at http://localhost:8765 (tree view + parallel branches); `writing-plans/scripts/ensure-dashboard.sh` or the `provledger-dashboard` slash command starts it.

Files in this skill

  • SKILL.md18.1 KB
  • reference/op-catalog.md9.1 KB
  • reference/when-to-use.md2.2 KB
  • scripts/_apply_op.py27.1 KB
  • scripts/_truncate_log.py1.6 KB
  • scripts/abandon-plan.sh992 B
  • scripts/agent-review-close.sh1.3 KB
  • scripts/append-log.sh1 KB
  • scripts/complete-step.sh1.1 KB
  • scripts/deviate.sh1 KB
  • scripts/fail-step.sh1 KB
  • scripts/finish-plan.sh1.1 KB
  • scripts/headline-respond.sh1.1 KB
  • scripts/raise-budget.sh1.1 KB
  • scripts/reason-fill.sh1.1 KB
  • scripts/reason-slots.sh1.1 KB
  • scripts/record-metric.sh1.1 KB
  • scripts/record-skill.sh1.1 KB
  • scripts/run-step.sh10.4 KB
  • scripts/start-step.sh1 KB

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…