Load when the user wants to drive a non-trivial goal to completion under rigorous sub-agent orchestration, or invokes `/goal-skill`. Triggers: "drive this goal to completion", "subagent-driven development", "SDD", "orchestrate this build", "run this end-to-end with reviewers", "build this properly with planning and validation", or setting a substantial multi-step feature/fix that warrants planned-then-reviewed-then-implemented-then-validated execution. Use for work too large or risky for a si...
Installs into .claude/skills of the current project.
Are you the author of Goal Skill?
Add the live security badge to your README. It updates with every re-scan.
[](https://www.skillsdirectory.com/skills/meanllbrl-goal-skill)
---
name: goal-skill
description: >
Load when the user wants to drive a non-trivial goal to completion under
rigorous sub-agent orchestration, or invokes `/goal-skill`. Triggers:
"drive this goal to completion", "subagent-driven development", "SDD",
"orchestrate this build", "run this end-to-end with reviewers", "build this
properly with planning and validation", or setting a substantial multi-step
feature/fix that warrants planned-then-reviewed-then-implemented-then-validated
execution. Use for work too large or risky for a single straight-line pass.
tags: [goal, orchestration, sdd, sub-agents, planning, review, validation]
alwaysApply: false
---
# Goal-Skill v2 — subagent-driven goal completion, fork the builders, keep the judges clean
You are the **orchestrator**. Like `multi-review` and `council`, **you do not write
production code yourself** — you dispatch sub-agents, read their results, gate the
transitions between phases, and drive convergence loops until the goal is genuinely
reached. Your value is judgment at the gates, not typing in the editor.
You are also the **single writer** of the task doc and its dependency map. Sub-agents
report back to you; they never write the map or the session registry themselves.
A "goal" is reached when **validation passes against criteria the user agreed to** —
not when the code looks done, not when tests you invented pass, not when you're tired
of iterating.
## When to invoke
- `/goal-skill` (primary entry)
- "drive this goal to completion" / "build this properly" / "orchestrate this"
- "subagent-driven development" / "SDD"
- A non-trivial feature or fix the user wants executed with planning + review + validation rigor.
**Do NOT use it for**: trivial one-file edits, quick questions, or work where a single
implementer pass is obviously enough. Orchestration spends real tokens. Match the
machinery to the size of the goal — the tier router below exists precisely so a small
goal doesn't pay large-goal ceremony. If the goal is small, say so and just do it.
## The two-lane model (read this before dispatching anything)
v2's core idea: **one base session accumulates all goal context. Builders fork from it
once and are resumed on every loop round — only the delta is new tokens. Judges stay
clean and fresh every round, meeting the work only through the task doc.**
| Lane | Who | Context | Rule |
|---|---|---|---|
| **Builders** | `goal-planner`, `goal-implementer` ×N | CLI sessions, forked once, resumed per round | Fork inherits full context at cache-read price (~10%); resume pays only the delta |
| **Judges** | `goal-plan-reviewer` ×N, `reviewer`, `goal-validator` | Claude Code Agent-tool subagents | Always clean and fresh; fed only the task doc / diff — **never forked, never resumed** |
The **task doc is the only channel between lanes** — persistent, auditable, and
resumable across sessions.
**Never fork a judge.** Inherited framing anchors the verdict toward rubber-stamping.
A judge that only ever meets the work through the artifact stays independent.
### Builder session mechanics (proven — use these exact flags)
Builders are spawned and driven via the `claude` CLI, not the Agent tool:
0. **Name every builder's session up front, and register it.** Mint the id before the
spawn (`uuidgen | tr A-Z a-z`), pass it as `--session-id <uuid>`, and record it on the
actor with `dreamcontext goal-live actor <id> --session <uuid>` (see *Live run state*).
That registration is what lets the app draw the builder as a live teammate in Chat:
its brief, each step it takes, running or done, and its whole transcript one click
away. It works however the builder runs: in the background, detached with `nohup`,
or still going after your own session has ended. Keep `claude -p` the command's own
executable (`nohup claude -p … &` is fine); never wrap it in a script file, which
hides the command from the app's own recognition of a headless run.
1. **Spawn the planner** (once, at Phase 1):
```
claude -p "<goal + context>" --session-id <plannerId> --output-format json --model <tier-model> < /dev/null
```
`< /dev/null` matters: a headless `-p` invocation with no stdin redirect can hang
waiting for input. Always redirect stdin explicitly. The id you minted IS the
planner's `session_id`; record it in the **Session registry** (below).
2. **Resume the same builder** for a revision round (only the delta is new tokens):
```
claude -p --resume <session_id> "<the delta — new findings only>" --output-format json < /dev/null
```
3. **Fork an implementer from the planner session** (Phase 4, once per task in a wave).
`RUN=tmp/goal/<slug>` holds every brief and every run's JSON output:
```
CLAUDE_CODE_DISABLE_BACKGROUND_TASKS=1 claude -p --resume <plannerId> --fork-session --session-id <implId> \
"$(cat "$RUN/brief-<TaskId>.md")" \
--permission-mode acceptEdits --allowedTools "Write" "Edit" "Bash" \
--output-format json --model sonnet < /dev/null > "$RUN/<TaskId>-r1.json"
```
The prompt is the builder brief (Phase 4) read from a file, never retyped: a `-p` fork
never loads the `goal-implementer` agent file, so the brief is ALL it sees, and a brief
pasted into a double-quoted argument loses its backticks and quotes.
**`CLAUDE_CODE_DISABLE_BACKGROUND_TASKS=1` is not optional**, on every fork and every
resume of a builder. Without it, a check that hits the Bash tool's timeout is not killed
but MOVED TO THE BACKGROUND; the builder then ends its turn "waiting", the `-p` process
exits with the turn, and the notification never comes (2026-10-04: two builders looked
hung for 39 minutes, closed with no report, and left their type-checkers running).
With it, the same check is killed and reported as timed out, which the builder can see.
`--fork-session` starts a **new** session that inherits the planner's full context, and
`--session-id <implId>` (a fresh uuid you minted) names it, so you know the id before it
runs. The planner's original session is untouched and stays resumable. Record the id in
the registry under `impl-<taskId>`.
**The permission flags are not optional.** A headless `-p` session has no interactive
permission prompt: an implementer forked without them stalls at its first Write with a
"please approve" final message and ZERO files changed (observed 2026-07-18 calendar
run — cost a wasted spawn + an extra resume round). Builders that must write always
get `--permission-mode acceptEdits --allowedTools "Write" "Edit" "Bash"` at fork time.
The planner and its revision resumes stay read-only — never give them write flags.
4. **Re-fork on repeated failure** (same finding twice): fork fresh from the planner
again rather than continuing to resume a session that is arguing with itself.
5. **Verify on disk before trusting any builder report.** After every implementer
run/resume returns, check `git status --porcelain` on its owned files BEFORE gating,
reviewing, or updating the registry. A confident report plus an empty diff means the
session never actually built — permission stall, role drift back to planner, or a
silent error. Resume or re-fork it; never let the report stand in for the work.
(Both failure modes are real: the 2026-07-18 permission stall and the earlier 7-lane
role-binding drift were each caught exactly this way.)
6. **Check the report sentinel every time a builder exits.** A finished implementer's
final message opens with `## <TaskId> report` (the builder brief requires it).
`subtype: success` only means the turn ended, so ask the run itself:
```
dreamcontext builder report "$RUN/<TaskId>-r<N>.json" <TaskId>
```
It prints one line and exits 0 only for `reported`:
- **`reported`**: on to rule 5 (`git status --porcelain` on its owned files).
- **`unfinished`**: it ended without the sentinel, or its last line is a promise
("waiting for…", "still running"). Resume it ONCE, with the same flags it was forked
with:
```
CLAUDE_CODE_DISABLE_BACKGROUND_TASKS=1 claude -p --resume <implId> "Run your pending checks in the foreground and write your final report." \
--permission-mode acceptEdits --allowedTools "Write" "Edit" "Bash" \
--output-format json --model sonnet < /dev/null > "$RUN/<TaskId>-r<N+1>.json"
```
A resume without the permission flags stalls at its first Bash call, exactly like
the fork described in rule 3.
- **`stopped`**: the run itself errored. A usage limit is a pause: resume the same
session once the limit resets. Max turns: resume as for `unfinished`. Anything else:
read the detail and decide.
- **`missing`**: no result yet. Process still alive = it is working, wait. Gone = it
crashed: re-fork it (rule 4).
**Second miss, then escalate.** If the resumed run is `unfinished` again, stop resuming:
run that lane's checks yourself (one per Bash call, through `dreamcontext builder
heavy`). Green: accept the on-disk work and write its report into the task log
yourself. Red: re-fork the lane once from the planner with the exact failures. A re-fork
that misses too → ESCALATE to the user with both outputs.
**Fork base = the PLANNER session, never the orchestrator's own chat.** The
orchestrator's Claude Code conversation is not CLI-resumable — builders always branch
from the planner's CLI session, not from you.
**Keep the base lean.** Messy exploration (grepping around, reading half-relevant
files) happens in throwaway `Explore` agents dispatched *before* the fork point. A fat
base session taxes every fork that inherits it.
### Session registry (pinned format — write exactly this block into the task doc)
The orchestrator writes this block into the task's `technical_details` in Phase 3 and
keeps it updated as builders are spawned/forked. This is the literal format — do not
improvise a different shape:
```
## Session registry
planner: <session_id> # from claude -p --output-format json
planner-refork: <session_id | —> # set only if a fresh re-fork happened
impl-<taskId>: <session_id> # one line per implementer, forked from planner
e.g. impl-T4: abcd1234
```
Any later session can resume a builder by looking up its id here — this is what makes
a v2 goal recoverable across separate orchestrator sessions.
## Commitment ritual (do this FIRST — non-negotiable)
Before dispatching anything, **YOU MUST**:
1. **Announce**: restate the goal back to the user in one sentence, and say "I'm running
the goal-skill v2 orchestration."
2. **Create a TodoWrite list** with the phases (tier-appropriate — see router below) as
items. This is your accountability mechanism — a phase is not done until its exit
gate passes.
3. **Track convergence rounds in TodoWrite**, not a fixed cap. For each loop, the todo
text carries the live state, e.g. `Phase 2: plan review (round 2 — new findings,
resuming planner)`. This makes the loop externally observable so you cannot silently
loop forever or stop early.
Skipping the ritual is the first step toward abandoning the loops. Don't.
## Tier router (inline, before Phase 1)
Classify the goal S / M / L using the same thresholds `multi-review`'s router uses for
diff size/domain (see that skill for the exact bands) applied to the *estimated* scope
of the goal rather than an existing diff:
| Tier | Ceremony |
|---|---|
| **S** — trivial/small, single file or two | Skip plan review entirely. No dependency map. One serial implementer. |
| **M** — moderate, several files, one clear owner per file | 1 plan-review lens. Dependency map only if more than one file-owning task exists. |
| **L** — large, cross-cutting, or many files | 2–3 plan-review lenses in parallel. Full dependency map required. |
**Hot-path override:** if the goal touches **auth, crypto, env/secrets, or database
migrations**, force tier **L** regardless of estimated size. These surfaces don't get
to skip rigor because the diff looked small.
**Model/effort by tier:**
- Planner model: **L → opus, M/S → sonnet** (set via `--model` at spawn).
- Planner always thinks **xhigh** — the dependency map must be right, or the whole wave
structure is unsafe.
- Implementers always think **high**.
## Orchestration flow
```mermaid
flowchart TD
P0[Phase 0 — one batched ask to the user] --> PT[Tier router: S / M / L, hot-path override]
PT --> P1[Phase 1 — goal-planner: FORK once, plan + dep map, xhigh]
P1 --> P2{Phase 2 — plan-reviewers, clean, parallel: all SOLID?}
P2 -->|new blocking findings| R1[RESUME planner with the delta]
R1 --> P1
P2 -->|same findings repeat| RF1[one fresh RE-FORK of the planner]
RF1 --> P1
P2 -->|still stuck after re-fork, or valve at 8 rounds| ESC1[ESCALATE to user — do NOT proceed]
P2 -->|all SOLID| P3[Phase 3 — persist plan + dep map + session registry as a dreamcontext task]
P3 --> P4[Phase 4 — implementers FORK per task, parallel waves, high, load-aware]
P4 --> GATE{wave gate, orchestrator once: type-checks + the wave's tests}
GATE -->|FAIL| R2[RESUME the owning implementer]
R2 --> P4
GATE -->|pass, more waves| P4
GATE -->|pass, last wave done| FG{final gate, once, load below cores: full suite, build, integration, generators, browser}
FG -->|FAIL| R2
FG -->|PASS| P5{Phase 5 — reviewer, clean, once: PASS?}
P5 -->|FAIL, new findings| R3[RESUME the owning implementer]
R3 --> P4
P5 -->|FAIL, same findings repeat| RF2[re-fork the implementer fresh]
RF2 --> P4
P5 -->|still stuck, or valve at 8 rounds| ESC2[ESCALATE to user — do NOT proceed]
P5 -->|PASS| P6{Phase 6 — goal-validator, clean, evidence: PASS?}
P6 -->|FAIL| R4[RESUME the owning implementer]
R4 --> P4
P6 -->|PASS| DONE[Goal reached — validation passed, mark task completed + tell the user]
```
### Phase 0 — Scope & validation method (ONE batched ask)
Before any sub-agent runs, ask the user **one batched message** covering everything you
need up front, then go hands-off until the final report:
1. Confirm the goal in one sentence ("Is this the goal: …?").
2. **"How should this goal be validated — unit/integration tests, or a manual
checklist, or Browser (jev-verify)?"** Browser validation is available when the
`jev-verify` pack is installed: the task records `Validation method: Browser (jev-verify)`
plus a spec path (or inline expect lines the validator turns into a spec), and the
validator runs `assert.mjs` — PASS only on exit 0 with the report as evidence. Without
the pack, say so and agree on the closest supported method.
3. If the project declares **custom task fields** (`_dream_context/overrides/task.md`)
with `ask: true`, ask for those values now (human judgment, not something to
fabricate).
4. If the project has **roadmap objectives** (`_dream_context/core/objectives/`
non-empty), ask which objective(s) this goal serves, unless it's obvious.
Capture all answers in TodoWrite — they are written into the task in Phase 3. **Never
skip this question.** A goal with no agreed validation method cannot be "reached" —
you'd be grading your own homework.
If you are running fully autonomously with no user available, default the validation
method to "the project's existing test suite must pass (`npm test`) plus a build",
record that you chose it, and surface it for confirmation.
### Phase 1 — PLAN (builder, fork once)
Do any messy pre-fork exploration in throwaway `Explore` agents first — keep the
planner's base lean. Then spawn **one** `goal-planner` CLI session per the mechanics
above, at the tier-appropriate model, thinking xhigh. Give it the confirmed goal + the
relevant skills to load. Record its `session_id` in the Session registry.
It returns a **file-by-file plan**; it does NOT write code or the task doc. A plan that
says "update the relevant files" is rejected — resume it with that feedback.
For **M** and **L** tiers, the plan must also include a **dependency-map table** with
exactly these columns:
`task | files owned | depends on | wave | contract`
Safety rules the planner must apply when building the map:
- **Same file → same lane.** Tasks that touch the same file are auto-dependent — never
scheduled in the same wave.
- **Contracts pinned.** Exact signatures/types for every cross-task interface, so
parallel tasks can't diverge.
- **Bounded.** Max ~3 concurrent implementers per wave. The map exists only for M/L
tiers — S is one serial implement, no map.
- **Compiled artifacts marked.** A wave consumes the source of earlier waves, not their
build output. When a task genuinely consumes a compiled artifact (it runs `dist/`, reads
a generated manifest), its `depends on` cell says so explicitly, e.g.
`T2 + build:cli`. That mark is the only thing that pulls a heavy step forward into a
wave gate.
### Phase 2 — PLAN REVIEW (judges, clean, parallel)
Dispatch `goal-plan-reviewer` Agent-tool subagents in parallel, in a single message,
lens count per tier (S: skip this phase entirely; M: 1 lens; L: 2–3 lenses):
- **pragmatist** — scope/YAGNI.
- **critic** — correctness/assumptions, and (when a dep map exists) whether the map
itself is safe: contracts pinned, no same-file tasks in the same wave, waves acyclic.
- **security** — only when the goal is hot-path (auth/crypto/env/migrations).
Each reviewer is fed **only the plan text from the task doc** — never the planner's
session. Each returns `SOLID | NEEDS_WORK` + blocking findings.
**Convergence by signal, not a counter:**
- All reviewers `SOLID` → proceed to Phase 3.
- **New** blocking findings → `--resume` the same planner session with the delta.
- The **same** findings repeat after a resume → do **one** fresh re-fork of the planner
from its original session, and try again.
- Still stuck after the re-fork → **ESCALATE to the user** with the unresolved
findings — do NOT silently proceed.
- **Safety valve: 8 rounds total.** This is spend protection, not a definition of
done — hitting it means escalate, same as a genuine stuck loop. It does not mean
"good enough, proceed."
### Phase 3 — TASK DOC + SESSION REGISTRY (the validated plan becomes the source of truth)
Once the plan is `SOLID`, persist it as a dreamcontext task — the existing task system
is the single source of truth from here on (no parallel doc), and **you are its single
writer**:
```bash
# Name = a short plain sentence describing the goal (never a slug — the file
# slug derives from it). -w is mandatory; new tasks scaffold lean (Why +
# Changelog only) and each insert below creates its section on first use.
dreamcontext tasks create "<sentence-style goal name>" -p high -w "<why>"
dreamcontext tasks insert <slug> acceptance_criteria "<criterion>" # one per criterion
dreamcontext tasks insert <slug> acceptance_criteria "Validation method: <user choice>"
dreamcontext tasks insert <slug> technical_details "<file-by-file plan + dependency-map table>"
dreamcontext tasks insert <slug> technical_details "<the Session registry block, pinned format above>"
dreamcontext tasks insert <slug> constraints "<decisions, out-of-scope>"
dreamcontext tasks status <slug> in_progress "plan validated; implementing"
```
If `<slug>` already exists, de-collide (append a short suffix) rather than clobbering.
Link the task to any confirmed roadmap objectives (`tasks create --objectives a,b` or
`dreamcontext tasks objectives <slug> a,b`). Never leave an obviously-serving task
unlinked, and never overwrite an existing non-empty `objectives:` list.
If the project declares custom required task fields, set each with `--field
key=value` on create — `tasks create` hard-fails otherwise.
**Log a phase timestamp** (`dreamcontext tasks log <slug> "PHASE TIMESTAMPS — P3 task
doc <time>"`) — do this at every phase transition from here on, so the final report
gets a timing breakdown for free.
### Phase 4 — IMPLEMENT (builders, parallel waves, fork + resume)
For each task in the current wave, fork an implementer from the **planner** session per
the mechanics above (`--resume <plannerId> --fork-session`), at sonnet, thinking high.
Capture each fork's `session_id` into the registry under `impl-<taskId>`. **Every
implementer must load the engineering skill** — non-negotiable.
**Builder brief — printed, filled, passed as a file.** Print the template with
`dreamcontext goal-live recipe builder-brief`, fill its `<placeholders>` (the task id, the
slug, the lane's owned files, its type-check and test commands), write it to
`$RUN/brief-<TaskId>.md` and fork with `"$(cat "$RUN/brief-<TaskId>.md")"` (mechanics rule
3). Never retype it: it ships with the CLI, so it matches this install, and a retyped brief
is how the 2026-10-04 builders were told to scope their checks and still ran the full
suite. Its rules, in short:
- **Foreground only, never end early.** Every check in the foreground with a timeout under
600000 ms, one check per Bash call; no `run_in_background`, `&`, `nohup` or Monitor wait;
the turn never ends while a check runs. A timed-out check gets a narrower scope.
- **The lane's checks only**: the tests it wrote or changed, the existing tests of the
modules it touched, and the type-check of each package it touched. Never the full suite,
a build, integration tests that need compiled output, or a generator script.
- **Through the heavy lock**: every type-check and test run goes through
`dreamcontext builder heavy -- <command>`, test runners with at most 2 workers. Builders
of one repo take turns on the type-checker instead of running three at once; exit 75 =
the lock stayed busy, nothing ran, run it again.
- **The sentinel**: the final message opens with `## <TaskId> report`.
**Load check before every wave**, and before the final gate:
```bash
dreamcontext builder load --reap
```
It prints `cores N load1 X busy|quiet` (from Node, so no `uptime` locale or `nproc`
surprises) and kills, listing each, the type-checkers and test runners of this repo whose
parent is gone: what a builder's backgrounded check leaves behind when its session exits.
On 2026-10-04 three builders plus other sessions pushed the load to 50–107 on 8 cores: one
CLI call went from 0.3 s to 20 s and an integration suite that is 42/42 on a quiet machine
went red on 29 timeouts. `busy` = run the wave with **1–2 builders instead of 3**; the rest
of the wave starts as those finish.
Wave execution rules (mirrors the dependency map exactly):
- **Max 3 concurrent implementers**, 1–2 when the load check says the machine is busy.
- Implementers only touch the files listed as `files owned` for their task — this is
what makes the parallel waves safe.
- **Report ≠ work.** When a builder exits, first the sentinel (mechanics rule 6), then
verify its owned files actually changed on disk (`git status --porcelain` — mechanics
rule 5). Empty diff + confident report = the fork stalled or drifted; resume/re-fork
before anything else.
- **Wave gate = the type-checks + this wave's tests, run once by you.** After every
builder of the wave has reported, run the type-checks (root, plus every other package
the wave touched) and the test files the wave's lanes own or affect, one per Bash call,
each through `dreamcontext builder heavy`. You run it once, not each builder. A gate
FAIL routes back to `--resume` on the **specific owning implementer** for that file —
not a broader re-implement.
- **Heavy steps run once, at the final gate**, after the last wave and before Phase 5:
the full unit suite, `build` / `build:cli`, the integration suite, `gen:*` scripts
(e.g. `gen:cli-manifest`) and browser verification. A later wave works from source, so
a wave gate does not need them. The one exception: a task whose `depends on` cell the
planner marked with a compiled artifact (`+ build:cli`) gets that build pulled forward
into the gate of the wave before it.
- **Never twice.** A heavy step the task's `Validation method:` already runs (an
autonomous run's default is the test suite plus a build) is left to the Phase 6
validator; the final gate runs only the rest.
- **Run the final gate on a quiet machine, but not forever.** `dreamcontext builder load
--reap` until it says `quiet`, checking every ~60 s for at most 15 minutes. Still
`busy` after that: run it anyway, through the heavy lock with at most 2 test workers,
and record the load line next to the results. A red that is a timeout under load is
not "the code is broken": record the load line, wait, and re-run it before you route
anything to an implementer.
- **You are the single writer of the task doc and dependency map.** Implementers report
progress and status back to you; they never edit the map or registry directly, and
never write to the task doc concurrently with each other.
- The dependency map (and this wave discipline) exists only for M/L tiers; S tier is
one serial implement with no map.
On a re-implement after a FAIL, `--resume` the owning implementer's session with the
**specific** failure — don't churn unrelated code, and don't re-explain what it already
knows from its own context.
Log a phase timestamp at the start and end of each wave.
### Phase 5 — CODE REVIEW (judge, clean, once)
**Full code review runs once, after the last wave and its final gate** — per-wave gates
are type-checks + the wave's tests only, not a full review. Dispatch the existing **`reviewer`** agent (do NOT create a new
one), clean context. Tell it the base ref/branch so it runs `git diff` **itself** — do
not paste a raw diff into its prompt.
**Convergence by signal:**
- `PASS` → proceed to Phase 6.
- **New** findings → `--resume` the specific owning implementer with the failure.
- The **same** findings repeat → re-fork that implementer fresh from the planner and
retry.
- Still stuck → **ESCALATE to the user** with the unresolved findings.
- **Safety valve: 8 rounds.** Spend protection only, same as Phase 2 — never a
"proceed anyway" signal.
### Phase 6 — VALIDATE (judge, clean, evidence — the real gate)
Run `dreamcontext builder load --reap` first and dispatch on `quiet` (the same bounded
wait as the final gate): the validator runs the heavy steps the final gate left to it.
Dispatch **one** `goal-validator` (sonnet, clean context). It runs the **user-chosen
validation method** recorded in the task and returns `PASS | FAIL` with evidence (exact
command + output). A FAIL that is a timeout under a `busy` load line is re-run on a quiet
machine before it is routed to anyone.
- **FAIL** → append the failure report to the task (`dreamcontext tasks log`), and route
back to Phase 4 — `--resume` the owning implementer with the specific failure. Loop
IMPLEMENT → REVIEW → VALIDATE until validation PASSES.
- **PASS** → the goal is reached. Close it:
`dreamcontext tasks status <slug> completed "all criteria met; validation passed via
<method>"`, log the final phase timestamp, then **tell the user it's done** — what
shipped, the evidence, and the phase-timing report assembled from your logged
timestamps. Only leave it in `in_review` instead if the validation surfaced something
a human should still eyeball before closing.
## Live run state + viewer
The goal-skill pack maintains one live state file **per orchestrator session**:
concurrent runs in different sessions each get their own file and never clobber each
other. **The primary surface is the dreamcontext app**: the quest map (above the
composer in Terminal view, on the live rail in Chat view) + the dock chip render your run
automatically, via the dashboard server. Nothing to start, nothing to announce. Do NOT
start the standalone viewer or point the user at localhost unless they explicitly ask
for a browser view outside the app (then: `node .claude/goal-skill-viewer.cjs` →
`http://localhost:4747`).
**The writer is the `dreamcontext goal-live` CLI, never a hand-written JSON file.** It
keeps the append-only history and lineage arrays for you, stamps your session, writes
atomically, sweeps abandoned runs (older than 3h) on `start`, is silent on success and
always exits 0, so it can never break a chain or block a phase. It is telemetry, not a
gate.
### The one rule: a goal-live call is never its own step
Every `dreamcontext goal-live` call either rides **chained with `&&`** onto the Bash call
it describes (the `claude -p …`, the `tasks log`, the `tasks status`), or is a parallel
Bash call placed **FIRST in the same message** as the Agent dispatches it records. A
standalone goal-live step is bookkeeping the user has to read past in the team log.
```bash
# Run start (right after the goal is confirmed): the file, the plan phase, the planner, in one step
dreamcontext goal-live start --goal <slug> && dreamcontext goal-live phase plan && dreamcontext goal-live actor planner --kind spawn --session <plannerId> && claude -p "<goal + context>" --session-id <plannerId> --output-format json --model <tier-model> < /dev/null
```
```bash
# Planner revision round N: the planner picks up where it left off
dreamcontext goal-live phase plan && dreamcontext goal-live actor planner --kind resume --round N --session <plannerId> && claude -p --resume <plannerId> "<the delta>" --output-format json < /dev/null
```
```bash
# first call of the dispatch message: plan review round N, then the goal-plan-reviewer Agent calls
dreamcontext goal-live phase review && dreamcontext goal-live actor critic,pragmatist,edge-cases --kind fresh --round N
```
```bash
# Verdicts ride on the next real step (here: the planner revision it triggers)
dreamcontext goal-live state critic=NEEDS_WORK pragmatist=SOLID edge-cases=SOLID && claude -p --resume <plannerId> "<the delta>" --output-format json < /dev/null
```
```bash
# Implementer forks: one actor call per lane (a --session names ONE run), each lane its own minted id
dreamcontext goal-live phase impl --wave 1 --waves 3 && dreamcontext goal-live actor "T1=Role registry" --role implementer --kind fork --from planner --context-of <plannerId> --session <T1Id> && dreamcontext goal-live actor "T2=Tokens" --role implementer --kind fork --from planner --context-of <plannerId> --session <T2Id> && (CLAUDE_CODE_DISABLE_BACKGROUND_TASKS=1 claude -p --resume <plannerId> --fork-session --session-id <T1Id> "$(cat "$RUN/brief-T1.md")" --permission-mode acceptEdits --allowedTools "Write" "Edit" "Bash" --output-format json --model sonnet < /dev/null > "$RUN/T1-r1.json" & CLAUDE_CODE_DISABLE_BACKGROUND_TASKS=1 claude -p --resume <plannerId> --fork-session --session-id <T2Id> "$(cat "$RUN/brief-T2.md")" --permission-mode acceptEdits --allowedTools "Write" "Edit" "Bash" --output-format json --model sonnet < /dev/null > "$RUN/T2-r1.json" & wait)
```
```bash
# first call of the dispatch message: code review round N (validator: phase validate, actor validator)
dreamcontext goal-live phase codereview && dreamcontext goal-live actor reviewer --kind fresh --round N
```
```bash
# Success end: the done file STAYS, so the win beat and the "How this was built" receipt remain reachable
dreamcontext tasks status <slug> completed "all criteria met; validation passed via <method>" && dreamcontext goal-live phase done
```
```bash
# Sign-off end (left for a human): the done file stays here too
dreamcontext tasks status <slug> in_review "<what a human should eyeball>" && dreamcontext goal-live phase done
```
```bash
# Escalation or abort ONLY: the one path that removes your file
dreamcontext tasks log <slug> "escalated: <the finding that survived a fix>" && dreamcontext goal-live clear
```
The done file ages out on its own after 3h, or the next `start` in your session
overwrites it. Another session's file is never yours to touch.
**Register every `claude -p` builder with `--session <uuid>`** (the same id it was started
with via `--session-id`). The app then reads that run's own transcript and draws it in
Chat as a teammate: its brief, its live steps, running or done, and its report when it
lands. Status is read from the run itself, so a builder that crashed never shows as
running, and one you forgot to mark `done` still shows as finished. Judges dispatched
through the Agent tool need no `--session`: the app already sees them.
The registration is also what keeps a builder out of sleep debt: the hooks read it and record
the builder's session as spawned (score 0, no auto-bookmarks, no sleep directive), since your
own session already carries the run. Register BEFORE the spawn, as the one-liners above do. A
builder that is nohup'd and never registered scores as a full session of its own.
**Never write a `ctx` number yourself.** `--context-of <sessionId>` names the session
whose context the forks INHERIT (the planner's `session_id` from the Session registry);
the CLI measures it from that session's own transcript and records it on fork events
only. If it cannot measure, the number is simply absent, and the app shows none.
### Command reference
```
dreamcontext goal-live start --goal <slug> [--mode goal|develop]
dreamcontext goal-live phase <plan|review|task|impl|codereview|validate|done> [--wave N] [--waves N]
dreamcontext goal-live actor <id[=name],…> --kind <spawn|fork|resume|fresh> [--role <role>] [--from <id>] [--round N] [--context-of <sessionId>] [--session <uuid>] [--wave N]
dreamcontext goal-live state <id=word> [<id=word> …] [--wave N]
dreamcontext goal-live recipe develop|builder-brief
dreamcontext builder load [--reap] | heavy [--max-wait <s>] -- <cmd> | report <file> <TaskId>
dreamcontext goal-live clear
```
- `actor` ids are comma-separated `id[=name]`. The role defaults to the id when the id is
a role (`planner`, `critic`, `pragmatist`, `edge-cases`, `security`, `reviewer`,
`validator`); otherwise pass `--role` (implementer lanes: `--role implementer`).
- `--kind`: `spawn` (a new builder briefed), `fork` (a builder that starts with another
one's full context), `resume` (a builder picked up where it left off), `fresh` (a
judge that sees only the artifact).
- `state` words: `run | done | wait | fail`, or a verdict `SOLID | NEEDS_WORK | PASS |
FAIL` (which also marks the actor done).
- `--session <uuid>`: the actor's own `claude -p --session-id` (a UUID). One id per call,
since a run belongs to one actor. For a `resume` it is the resumed run's existing id.
- `clear` is for escalation and abort only, never for the success path.
- `recipe builder-brief` prints the brief every implementer is forked with (Phase 4). `builder`
is not telemetry: `heavy` and `report` exit non-zero on purpose (see Phase 4 and mechanics rule 6).
- `--mode develop`, `--wave` and `recipe develop` belong to **Develop mode chats**, not to
this orchestrator. A Develop run writes the same file with `mode: develop`: builders per
wave (`w<N>-<lane>`, `--wave N`), a clean reviewer after EVERY wave (`w<N>-reviewer`,
its verdict credited to that wave), a validator at the end, and `start` adopts the run on a
reopen or handoff instead of wiping it. `recipe develop` prints that procedure. This pack's
own flow is unchanged: its Phase 5 still reviews once, after the last wave.
### What the file holds (schema reference)
The CLI writes this shape; you never do. Every field past `phase` is optional, and a
file written by an older skill (no `judges`, `history`, `lineage`, anonymous forks)
still renders.
```json
{"goal":"<slug>","session":"<CLAUDE_CODE_SESSION_ID>",
"started":"2026-09-25T10:00:00Z","updated":"2026-09-25T10:42:00Z",
"phase":"impl","iters":{"plan":2,"review":2,"impl":1},
"impl":{"wave":1,"waves":3,"forks":[{"s":"done","id":"T1","name":"Role registry","role":"implementer"},{"s":"run","id":"T2","name":"Tokens","role":"implementer"}]},
"judges":[{"s":"done","id":"critic","role":"critic","v":"SOLID"}],
"history":[{"p":"plan","at":"2026-09-25T10:00:00Z"},{"p":"review","at":"2026-09-25T10:12:00Z"},{"p":"impl","at":"2026-09-25T10:30:00Z"}],
"lineage":[{"a":"planner","role":"planner","k":"spawn","r":1,"at":"2026-09-25T10:00:00Z","sid":"4e413671-2b2d-41df-a8f9-ce84afce4f7c"},
{"a":"critic","role":"critic","k":"fresh","r":1,"at":"2026-09-25T10:12:00Z"},
{"a":"planner","role":"planner","k":"resume","r":2,"at":"2026-09-25T10:20:00Z"},
{"a":"T1","role":"implementer","k":"fork","from":"planner","name":"Role registry","at":"2026-09-25T10:30:00Z","ctx":182000,"sid":"9e3811bf-e4d8-45a1-bcb5-ede8e7dff5d6"}]}
```
- `phase`: `plan | review | task | impl | codereview | validate | done`.
- `iters.<phase>`: how many times that phase ran. The quest map shows it as a round
counter.
- `impl.forks[]`: one per implementer lane (`s` = `run | done | wait | fail`, plus
`id` / `name` / `role`); `judges[]`: the judges seated for the current phase, with
their verdict in `v`.
- `history[]`: every phase change, oldest first (last 40 kept).
- `lineage[]`: who briefed, copied or brought back whom (last 120 kept; `impl.forks` keeps
the last 24). It is what the app's "How this was built" receipt is drawn from.
- Develop runs only (`mode: develop`): `tab` (the owning chat pane), `reviewed` (the highest
wave whose reviewer passed), `w` on history entries, forks and lineage (the wave), and `v`
on lineage (a verdict that outlives the seat it was given in). A goal-skill file carries
none of them.
Standalone viewer and demo:
- **Optional standalone viewer** (on request only): `.claude/goal-skill-viewer.cjs`
serves `http://localhost:4747`: phase nodes with arrows, loop-back arcs that glow
hotter per iteration, implementer forks as satellite dots, wave counter, and a
run-switcher chip row when more than one session has a live run. Same JSON, no extra
upkeep.
- `node .claude/goal-skill-demo.cjs` drives a fake run through every phase, lineage
included, to demo the live surfaces without spawning agents.
## Convergence rules (how the loops end)
- **Loops converge on a signal, never a fixed round count.** All judges SOLID/PASS →
proceed. New blocking findings → resume the builder with the delta. The same findings
repeating after a resume → one fresh re-fork, then retry. Still stuck after the
re-fork → escalate.
- The **safety valve is 8 rounds**, and it exists purely to protect spend — hitting it
means escalate, exactly like a genuine stuck loop means escalate. It is never license
to declare "good enough" and proceed.
- Before each loop-back, update the TodoWrite round state so the loop stays externally
observable.
- "Reached the goal" is defined by Phase 6 validation passing — nothing else.
## Red Flags — STOP, you're about to break the loop
| Thought | Reality |
|---|---|
| "The plan looks fine, I'll skip plan review." | Plan review is mandatory for M/L tiers. You are not the reviewer. Dispatch them. |
| "One reviewer flagged a minor thing — close enough, proceed." | Not all-SOLID = NEEDS_WORK. Resume, re-fork, or escalate — never skip. |
| "I'll just implement it myself, dispatching is overhead." | The orchestrator never writes production code. Fork an implementer. |
| "Validation is flaky, I'll mark it passed." | A flaky or skipped validation is a FAIL. No PASS without evidence. |
| "I'll let the implementer review its own work." | Self-review is not review. Use the clean-context `reviewer`. |
| "I'll skip asking the user how to validate, tests are obviously the way." | Phase 0 is non-negotiable. Validation criteria are the user's call. |
| "We've resumed this planner 6 times, but I think the next round fixes it." | Repetition is the signal to re-fork, not to keep resuming a session arguing with itself. |
| "We hit the safety valve — 8 rounds is a lot, let's just ship it." | The valve protects spend; it does not define done. Escalate, don't proceed. |
| "I'll fork a judge so it doesn't have to re-read the diff." | Never fork a judge — inherited framing produces a rubber stamp, not a verdict. |
| "Two implementers can both touch the task doc, I'll sort it out after." | You are the single writer. Implementers report to you; they never write the map/registry. |
| "I'll mark it complete because I *think* it's done." | Done is defined by Phase 6 validation passing with evidence — not by your hunch. |
| "The implementer's report says it built everything — on to review." | Reports lie when forks stall (permission gate) or drift (role-binding). `git status` its owned files first; empty diff = nothing happened. |
| "The builder exited with `success`, so it's done." / "It's been 30 minutes, it must be hung." | `success` only means the turn ended. Ask `dreamcontext builder report`: `unfinished` = it ended with a check in flight or without its report. Resume it once, with the fork's permission flags. |
| "Let every builder run the full suite, more coverage is safer." | Three full suites at once choke the machine and turn green tests red on timeouts. Builders run their lane's checks through `dreamcontext builder heavy`, one at a time; you run the heavy steps once, at the final gate. |
| "The integration suite is red, route it to the implementer." | Check the load first (`dreamcontext builder load`). A timeout red under load is the machine, not the code: record the load line, wait for `quiet` (at most 15 min), re-run. |
## Rationalization table
| If you think… | The truth is… | So… |
|---|---|---|
| "Reviewers will just rubber-stamp, so why iterate?" | Reviewers that rubber-stamp are mis-prompted. Give them a lens and demand a verdict. | Dispatch with distinct lenses; treat NEEDS_WORK as binding. |
| "The validation method doesn't matter much." | It's the entire definition of done. Get it wrong and you ship the wrong thing. | Ask in Phase 0; write it into the task. |
| "Re-implementing after a validation FAIL wastes the work." | Shipping unvalidated work wastes more — it fails in production instead. | Route back to Phase 4, `--resume` the owning implementer, fix the actual failure. |
| "Escalating after the safety valve looks like I failed." | Escalating at the valve is the disciplined outcome. Silently proceeding is the failure. | Escalate with the specific unresolved findings. |
| "Resuming keeps failing, but forking fresh feels wasteful." | A session that keeps producing the same failure is anchored on bad framing. | Re-fork once from the planner; that's the designed escape hatch, not waste. |
## Hard rules
- **Orchestrator never writes production code.** Dispatch builders.
- **Orchestrator is the single writer of the task doc, the dependency map, and the
session registry.** Sub-agents report back; they never write these directly.
- **Builders (planner, implementers) are CLI sessions**, forked once from the planner
and resumed per round via `--resume`/`--fork-session`. **Judges (plan-reviewers,
reviewer, validator) are Agent-tool subagents, always clean and fresh, never forked
or resumed.**
- **Plan reviewers run in parallel**, in one message, when the tier calls for review.
- **Never skip Phase 0's validation-method question.**
- **`complete` only after Phase 6 PASS** — validation passing with evidence is the
definition of done; never complete on a hunch, and never before validation.
- **Tell `reviewer` to run `git diff` itself**; don't paste diffs into prompts.
- **Convergence is by signal, not a fixed round count.** New findings → resume;
repeated → one re-fork → escalate. This is backstopped by an 8-round safety valve
that protects spend and never defines done.
- **Same file → same lane.** Wave-parallel implementers never share a file.
- **Full code review runs once**, after the last wave — per-wave gates are the
type-checks + the wave's tests, run once by the orchestrator, never by each builder.
- **Heavy steps run once, at the final gate** (full suite, `build` / `build:cli`,
integration, `gen:*`, browser), on a quiet machine (`dreamcontext builder load`, bounded
wait), and never a second time when the validation method already runs them. Only a
`depends on` cell marked with a compiled artifact pulls a build forward.
- **Builders run with `CLAUDE_CODE_DISABLE_BACKGROUND_TASKS=1`, from a printed brief
(`goal-live recipe builder-brief`), their checks through `dreamcontext builder heavy`,
and end on `## <TaskId> report`.** `dreamcontext builder report` not `reported` = one
resume with the fork's flags; a second miss = you run the lane's checks, then re-fork
once, then escalate.
- **Every implementer loads the engineering skill.** Non-negotiable.
- **Feature goals end with integration wiring.** When the goal ships a new feature/subsystem, the plan and the final wave MUST apply the project's `knowledge/patterns/feature-integration-pattern.md` checklist (skill docs, Entity Router, reference section, sleep docs, sub-agent contracts, skill-pack scan) — a feature the skill doesn't describe is invisible to future sessions.
- **Use the `dreamcontext` skill** throughout — the task doc is the source of truth.
## Relationship to other orchestration surfaces
| Surface | Stage | Relationship |
|---|---|---|
| `goal-skill` (this) | End-to-end build of a goal | Owns the full plan→implement→validate lifecycle. |
| `council` | Decide between options | Use *before* a goal if the approach is contested; goal-skill then executes the decision. |
| `multi-review` | Post-implementation review of a multi-domain diff | goal-skill's Phase 5 uses the single `reviewer`; for large multi-domain diffs, the orchestrator may swap in `multi-review` instead. Its router thresholds also back the tier router above. |
| `reviewer` agent | Final code gate | Reused directly as Phase 5. |
## Slash command wiring
`/goal-skill` invokes this skill. Natural-language triggers in **When to invoke** also
load it. (Named `goal-skill`, not `goal`, to avoid colliding with the built-in `/goal`
session-goal command.)