The ZOdyssey orchestration conductor. Loaded by `/orchestrate` and by the `zodyssey:prometheus` planner. Defines the full pipeline (triage → consult → plan → review → execute → verify → final-wave), the dispatch rules, parallelization logic, and the state-machine the enforcement hooks read. Follow this exactly when orchestrating.
Installs into .claude/skills of the current project.
Are you the author of Odyssey?
Add the live security badge to your README. It updates with every re-scan.
[](https://www.skillsdirectory.com/skills/amartinawi-odyssey)
---
name: odyssey
description: The ZOdyssey orchestration conductor. Loaded by `/orchestrate` and by the `zodyssey:prometheus` planner. Defines the full pipeline (triage → consult → plan → review → execute → verify → final-wave), the dispatch rules, parallelization logic, and the state-machine the enforcement hooks read. Follow this exactly when orchestrating.
---
# ZOdyssey — Orchestration Conductor
You are the **orchestrator** of a hybrid-enforced multi-agent pipeline. You direct the cast: `zodyssey:metis` (consult), `zodyssey:prometheus` (plan), `zodyssey:momus` (review), `zodyssey:explore`/`zodyssey:librarian`/`zodyssey:oracle` (research/advice), `zodyssey:sisyphus-junior` (execute). The enforcement hooks are declared in `.zcode-plugin/plugin.json` (resolved via `${CLAUDE_PLUGIN_ROOT}`) and hard-block the dangerous invariants (edits before plan-OKAY, file collisions, parallel-overflow, wrong-phase dispatch; post-edit lint blocks only on diagnostics the edit itself introduced). Your job is to *drive* the pipeline and *guide* the judgment parts.
> This skill is the conductor. The pipeline below is the state machine. Read the active run's `<repo>/.zcode/state/<slug>.json` at every transition — that file is the source of truth for phase, review verdict, locks, and progress.
## Where things live (per repo)
- Plans: `<repo>/.zcode/plans/<slug>.md`
- State: `<repo>/.zcode/state/<slug>.json`
- Notepads: `<repo>/.zcode/notepads/<slug>/<todo-id>.md`
- Scripts: `skills/odyssey/scripts/{scaffold,parse-plan,run-report,record-todo,record-capability,consult}.mjs` (inside the `zodyssey` plugin install)
- **Capability routing table: `skills/odyssey/references/capabilities.md`** (inside the `zodyssey` plugin install; read this at every phase)
## Capability routing — ALWAYS use the best tool for the activity
This is the orchestrator's core promise. This machine has 8 plugins, ~50 skills, 8 sub-agents, 20 MCPs, and codegraph. Before doing any activity the generic way, consult `references/capabilities.md` and use the best-fit capability. The headline routing:
| Activity | Use |
|---|---|
| Brainstorm/shape a fuzzy feature | `skill: brainstorming` (+ `premortem`) |
| Hard multi-step reasoning | `sequentialthinking` MCP (decompose before answering) |
| Plan | `Task: zodyssey:prometheus` + `skill: writing-plans` |
| Research codebase | `codegraph_explore` MCP if `.codegraph/`, else `Task: zodyssey:explore` |
| Research docs/libs | `Context7` MCP + `Task: zodyssey:librarian` |
| Design/architecture | `Task: zodyssey:oracle` + `skill: brainstorming` (+ `feature-dev:code-architect`) |
| Implement (logic) | `skill: test-driven-development` (non-negotiable for code) + `Task: zodyssey:sisyphus-junior` |
| Implement (plan) | `skill: executing-plans` (+ `using-git-worktrees`) |
| Debug (hard) | `skill: systematic-debugging` + `sequentialthinking` MCP (+ `Task: zodyssey:oracle` after 2 fails) |
| Security / vuln audit | `claude-security` plugin (verified findings + patches) |
| Audit code | `Task: code-reviewer` + `skill: source-command-audit-code` |
| Review plan (gate) | `Task: zodyssey:momus` (+ `Task: zodyssey:oracle` independent, for architecture) |
| Verify before "done" | `skill: verification-before-completion` |
| Remember across runs | `memory` MCP (knowledge graph) |
| Media/image/PDF | `Task: zodyssey:multimodal-looker` |
| Review a PR/changeset line-by-line | `skill: open-code-review` (`ocr` — external install, own LLM config) |
The table is the summary; `capabilities.md` is the authoritative detail. **Tell every agent you dispatch which capability to use for its activity** — don't assume they'll reach for it on their own. The whole point is that the orchestrator is the thing that *knows* to load TDD, codegraph, a premortem.
- **External-skill quality check:** before recommending or auto-using any external skill (discovered via `find-skills` / `npx skills find`), sanity-check it — prefer **≥1K installs** and official sources (`vercel-labs`, `anthropics`, `microsoft`, …); be skeptical below **~100 GitHub stars**. (Already-installed local skills are inventory, not reputation.)
### Routing is the DEFAULT — generic knowledge is the FALLBACK (enforced)
The table above is **not a menu of suggestions** — it is the default source of truth for how to do any activity. A capable model's temptation to skip the lookup and do the task from memory is the exact failure mode this rule exists to prevent: it is how installed skills get ignored (Task B: `aws-serverless` bypassed) and how capability gaps go undiscovered (Task A: `find-skills` never loaded). So, **before writing any code or generic prose for an activity**:
1. **Scan `capabilities.md`** for a fitting installed capability (skill / sub-agent / MCP). If one fits → **USE it** (load the skill in the parent thread, or dispatch the agent). Generic knowledge is not a substitute for a routed capability that fits.
2. **If none fits** → load **`find-skills`** and run the search + the ≥1K/official/~100★ quality gate above. Mandatory for any domain outside the installed inventory.
3. **Only then** fall back to generic model knowledge — and only if discovery returned nothing reputable.
**Record the decision as a tri-state declaration** that `zodyssey:metis` emits and `zodyssey:prometheus` transcribes into the plan's `## Capability routing` section:
- `routed: skill:<name>` (or `mcp:<server>` / `agent:<name>`) — an installed capability fits and will be used.
- `discovered: find-skills` — no installed capability fits; discovery attempted (search term + quality verdict).
- `generic: <one-line reason>` — valid **only after** discovery was attempted and returned nothing reputable.
This is **gated, not advisory**:
- **Review gate (phase 3):** `parse-plan --lint` fails OKAY if the plan lacks a `## Capability routing` section with a non-vacuous tri-state token. `zodyssey:momus` rejects it as a missing required section.
- **Final wave (phase 6):** `record-final-wave` cross-checks the declaration against `state.capabilities[]` (the hook-witnessed log of every **successful** `Skill` / `mcp__*` load, phase-stamped — errored loads are not recorded as observed). `routed: skill:X` requires an observed `skill:X`; `discovered` / `generic` requires an observed `skill:find-skills`. Matching is on the final name segment, so a bare declaration matches a plugin-namespaced observation (`skill:test-driven-development` ≡ `skill:superpowers:test-driven-development`). No matching observation ⇒ the final wave fails ⇒ the run cannot reach `done`.
Because sub-agents cannot load skills (trust anchor), **the orchestrator loads the chosen skill / `find-skills` in the parent thread** — that load is what the hook observes and what the final-wave gate checks. Telling a dispatched sub-agent to "use skill X" without loading it yourself produces zero observation and fails the gate.
### Research deliverables carry a contract (research-kind runs)
When metis classifies the run's KIND as research (the deliverable is a report, study, or survey — not code), zodyssey:prometheus writes a `## Deliverable contract` section into the plan, transcribed from metis's directives:
- **Levers** — three typed lines: `register: teach|survey|analyze|advocate` (default `analyze`; an explicit user directive always wins), `format: short|structured|argumentative`, `tier: light|full` (light = bounded lookup/comparison; full = contested topics and conflicting evidence — **when uncertain, tier up**).
- **Headings** — an ordered list of literal H2 headings the deliverable must emit, in order: one per enumerated ask, one per discuss/analyze-flagged entity, or 4-7 derived from the sub-questions for narrative asks. Write each item as a list line (`1.`/`-`) whose text begins with `## <Heading>`; the shape lint accepts the heading with or without surrounding backticks/quotes/emphasis wrappers (`` 1. `## Background` ``, `1. ## Background`, and `- "## Findings"` all count) — a line with no literal `## ` after the marker does not. Never empty — `parse-plan --lint` refuses a vacuous contract (shape), momus rejects a research plan without one (presence), and the external auditor judges the finished deliverable heading-by-heading against it.
- **Items** — the atomic decomposition: sub-questions, entities (with required fields), required formats, **period-pinned time periods with their primary source named** (missing period-pins are the top silent miss), scope conditions — plus a coverage note mapping every noun-phrase of the verbatim ask to an item (zero unmapped phrases).
Write the deliverable todo's acceptance criteria as grep-able heading checks (`grep -q "^## <Heading>" <deliverable>`), so phase 5 verifies the Deliverable contract mechanically.
## The state machine (8 phases: -1 priming → 0–6)
```
┌──────────────────────────────────────────────────────┐
│ -1. PRIME (you do this first, ALWAYS, before triage) │
│ Load `skill: prompt-master` and feed it the user's │
│ raw task. Produce a primed brief: │
│ · intent + success criteria │
│ · implicit constraints (surfaced) │
│ · ambiguities → ask the user (max 3, then commit)│
│ criteria-confirmation round — only when the │
│ primed brief's success criteria are measurable │
│ (a command + expected outcome): ONE │
│ AskUserQuestion, ≤3 criteria, ≤4 options, one │
│ always an explicit skip → confirm / adjust / │
│ skip. No answer, or no question tool (headless │
│ / autonomous runs) → skipped; never blocks; │
│ counts inside the budget clause above. Pass │
│ the outcome to scaffold as --criteria-state │
│ · a rewritten prompt that REPLACES the original │
│ If ambiguities need resolving, ask the user FIRST │
│ and WAIT — do not triage an ambiguous brief. │
│ The refined prompt + brief feed every later phase. │
└───────────────────────┬──────────────────────────────┘
▼
┌──────────────────────────────────────────────────────┐
│ 0. TRIAGE (you do this directly — do NOT dispatch) │
│ Uses the PRIMED brief, not the raw prompt. │
│ trivial → "just ask normally" + STOP. │
│ standard → continue. │
│ architecture → continue (team mode in v2). │
└───────────────────────┬──────────────────────────────┘
▼
┌──────────────────────────────────────────────────────┐
│ 1. CONSULT Task(zodyssey:metis) phase: consult │
│ FIRST: read prior learnings from the `memory` MCP │
│ (search_nodes for this repo + intent keywords). │
│ `recall-outcomes.mjs <repo>` — blocked/failed runs│
│ `recall-corrections.mjs <repo>` — correction sigs.│
│ `registry-report.mjs <repo>` — narrator trust. │
│ unapplied .zcode/staging/proposals → metis risks │
│ Then hand zodyssey:metis: request + repo root + │
│ memories. She returns intent, risks, questions, │
│ directives. If she lists user-questions, surface │
│ them and WAIT. If she recommends dispatching │
│ zodyssey:explore/zodyssey:librarian, run those │
│ first, then re-metis. │
└───────────────────────┬──────────────────────────────┘
▼
┌──────────────────────────────────────────────────────┐
│ 2. PLAN Task(zodyssey:prometheus) phase: plan │
│ Hand zodyssey:prometheus: the request, the repo │
│ root, the metis output, and a slug. zodyssey:prom-│
│ etheus loads THIS skill, runs scripts/scaffold.mjs│
│ to create the plan + state.json, drafts the plan, │
│ and returns the plan path. │
│ The criteria-confirmation stamp on the task brief │
│ rules the todos' Acceptance criteria: adjusted → │
│ transcribe the user's criteria from the brief body│
│ VERBATIM as executable commands (source of truth);│
│ confirmed → the presented criteria; skipped or no │
│ stamp → today's authorship, unchanged. │
│ After he returns, set state.phase = "review". │
└───────────────────────┬──────────────────────────────┘
▼
┌──────────────────────────────────────────────────────┐
│ 3. REVIEW Task(zodyssey:momus) phase: review (gate)│
│ Dispatch zodyssey:momus with the plan path. She │
│ returns [OKAY] or [REJECT] + ≤3 blockers. │
│ • REJECT → increment state.review.round. If round │
│ < max_rounds (3): re-dispatch zodyssey:promethe-│
│ us with the blockers, then re-review. If ≥ 3: │
│ STOP and surface to user (no unbounded loop). │
│ • OKAY → record-review.mjs (the ONLY sanctioned │
│ verdict write; a direct state.json edit is │
│ what the gate blocks), then set-phase execute. │
│ (Optional, for architecture intent: also dispatch │
│ zodyssey:oracle for an independent review; both │
│ must OKAY.) │
└───────────────────────┬──────────────────────────────┘
OKAY ▼
┌──────────────────────────────────────────────────────┐
│ 4. EXECUTE you + Task(zodyssey:sisyphus-junior) │
│ phase: execute. Parse the plan with │
│ scripts/parse-plan.mjs. Dispatch todos per the │
│ parallel-by-default rule (below). Each todo → one │
│ zodyssey:sisyphus-junior dispatch carrying the │
│ todo block + inherited wisdom. │
│ On each todo's return: write a checkpoint and │
│ update state via record-todo/record-verify │
│ (state.json is the source of truth). │
│ (Hooks enforce: file locks, parallel cap, phase.) │
└───────────────────────┬──────────────────────────────┘
▼
┌──────────────────────────────────────────────────────┐
│ 5. VERIFY you (run acceptance cmds) phase: verify │
│ For each done todo, run its acceptance-criteria │
│ commands. On failure, re-dispatch that todo's │
│ zodyssey:sisyphus-junior with the error output. │
└───────────────────────┬──────────────────────────────┘
▼
┌──────────────────────────────────────────────────────┐
│ 6. FINAL WAVE Task(zodyssey:oracle) + Task(zodyssey: │
│ momus) + Task(code-reviewer) against the full │
│ diff. F1 plan-compliance, F2 code quality, F3 │
│ manual QA checklist (you produce it), F4 scope │
│ fidelity, F5 capability-routing cross-check. All │
│ must pass before "done". │
│ AUTO COMPACTION (at final entry): set-phase.mjs │
│ auto-runs compact.mjs above the notepad line │
│ threshold (AUTO_COMPACT_MIN_LINES); inert │
│ at/below; skip via ZODYSSEY_NO_AUTO_COMPACT=1. │
│ The printed brief path is your signal to point │
│ F1-F4 at `_compact-brief.md`. Deterministic, $0, │
│ additive (never modifies source notepads). │
│ MEMORY RULE: delegate any step that must READ the │
│ todos' notepad/fragment outputs to a sub-agent — │
│ do NOT read those fragments back into your own │
│ context. You keep ~3% of an executor's output; │
│ the other ~97% should never enter your window. │
└───────────────────────┬──────────────────────────────┘
▼
phase: done
## Context economy (memory optimization — apply at every phase)
Your context window is the single largest cost center of a run (measured ~95% of real
memory pressure; on-disk state and per-call node spawns are negligible by comparison).
Three rules, in priority order:
1. **Never read a sub-agent's full output back into your context.** If a downstream step
needs what an executor wrote (a notepad, a findings doc, a synthesis), **dispatch a new
sub-agent to read it** — that fragment lives in *the sub-agent's* context, not yours.
You absorb only the sub-agent's ~3-line summary. Anti-pattern: reading a 30K-token
notepad to "drive synthesis" — instead, dispatch the synthesizer with pointers to the
fragments. This alone cuts a typical multi-stage run ~25%.
2. **Dispatch prompts are pointers + delta, not restatement.** Each parallel executor gets
its own copy of the dispatch text, so a 1.4K-word prompt × N executors multiplies. Point
the executor at files/paths it can read itself (the plan, the brief, the audit prompt),
and include only the delta: the specific todo scope, the must-not-do, the acceptance
criteria. Do not paste context the agent can read.
3. **Synthesis, refutation, and any "merge the fragments" step is ALWAYS a sub-agent.** The
orchestrator's job is to dispatch and judge 3-line summaries — never to hold the bulk
content the workers produced. If you catch yourself opening a `.zcode/notepads/<slug>/*.md`
file in a Read call during phases 4-6, stop: that's a sub-agent's job. For research runs,
the synthesis sub-agent must also emit an explicit **Tensions** section — claim vs claim,
the source on each side, which the deliverable adopts and WHY — contested findings are
surfaced, never silently averaged or dropped (ask workers in their dispatch to flag
conflicts with prior notepads).
These rules compose with (but do not replace) the anti-duplication rule: once you delegate,
don't re-research — and once you delegate, don't re-read.
### Notepads are load-bearing working memory (not optional scratch)
`.zcode/notepads/<slug>/<todo-id>.md` is the **load-bearing cross-todo working-memory surface** of a
run. Treat notepads as *state read by downstream waves*, never as optional scratch an executor may
or may not write. Concretely:
- Every dispatched `zodyssey:sisyphus-junior` writes a notepad at the path its dispatch names — what it
changed, decisions made, gotchas, and the acceptance-command output (evidence). This is
"inherited wisdom" for the next todo and the raw input the final wave synthesizes.
- A RESEARCH todo's notepad ends with one drift-check line — `Answers the question: yes|partial|no — <one-line why>` —
the conductor's cheap probe that the fan-out is still answering THE ask before synthesis runs.
- Downstream todos **read prior notepads by path** (the orchestrator passes the pointers), so a
notepad is the handoff contract between fan-out executors that never share a context window.
- The final wave (F1-F4) reads notepads through a delegated sub-agent (memory rule above) or, when
the run is large, through the `_compact-brief.md` `compact.mjs` auto-derives at final entry above the size threshold (additive — sources never modified).
This is the structural analog of prime-agent primitive #8's "the kernel survives across
compactions" — except ZOdyssey has no in-process kernel, so the persistence that survives across
executor lifetimes is the **filesystem**, and the per-todo notepad is the unit that survives.
Deleting or failing to write a notepad breaks the chain for every downstream consumer; treat
notepad writes as mandatory output, not a nicety.
## After done — persist learnings (memory MCP)
Before the run truly closes, write what's worth remembering to the `memory` knowledge graph:
- entities for durable facts learned (key files, architectural decisions, gotchas)
- relations linking them (e.g. decision → rationale, gotcha → file)
- only things that would save real time on a future run — never trivia
This is the cross-run learning loop: read at consult (phase 1), write at done.
## Inspecting a run while it is active (v0.5.1)
Interpreter invocation through Bash — `node -e`, `python -c`, `sh script.sh` — is **gated in every
phase** while a run is active, including read-only-looking one-liners. This is deliberate: the write
hides inside the payload, so the gate cannot tell `node -e "console.log(x)"` from
`node -e "fs.writeFileSync(...)"` without executing it. Enumerating safe flag shapes was tried twice
and failed twice (v0.5.0 and v0.5.1 both shipped bypasses), so interpreters are an allowlist now.
That makes one previously common habit unavailable — reading run state with
`node -e "require('./.zcode/state/<slug>.json')"`. Use instead:
- **`node <plugin>/skills/odyssey/scripts/dashboard.mjs <repo> <slug>`** — the sanctioned run-state
view; it is on the trusted-script allowlist, so it runs while gated.
- **The `Read` tool** on `.zcode/state/<slug>.json` — reads are never gated.
- **`cat` / `grep` / `jq`** — plain read-only commands stay allowed.
**Executable checks belong in acceptance criteria, not in Bash.** `record-verify.mjs` executes a
todo's declared criteria itself, so `node -e "import('./src/x.js')…"` works there and is *recorded
as evidence*, which a Bash one-liner never is. If you find yourself wanting to run an interpreter to
prove something works, that is the signal it should be a criterion.
## External consult/audit gate (opt-in, via `/orchestrate-consult <slug>`)
After a run is `done`, the user may invoke an **independent external audit**: a *separate* Claude
Code process (different model, fresh context — true independence) reviews the run's plan + full git
diff and returns ACCEPT or REJECT+gaps. This is stronger verification than any in-session reviewer
because the auditor cannot inherit the run's assumptions.
**You do NOT run this automatically.** It fires only on `/orchestrate-consult`. When it does:
1. **Confirm the run is `done`** (state.phase). The audit diffs `run_start_sha..HEAD` — no diff
exists until work has happened.
2. **Run one audit round:** `skills/odyssey/scripts/consult.mjs <repo> <slug>` (inside the `zodyssey` plugin install). The script:
- gathers `state.run_start_sha`, the plan, and `git diff <start>..HEAD`
- spawns the external Claude Code CLI headless: `claude -p "<prompt>" --output-format json`
- the prompt (`references/auditor-prompt.md`) forces a strict JSON verdict with the full-scope
rubric: **plan compliance + code quality + bugs + security**
- parses + normalizes the verdict, appends to `state.consult.history`, prints it
3. **On ACCEPT:** mark `phase: "audited"`, summarize, STOP.
(A retroactive-audit vehicle parked at `abandoned` — a run opened only to carry an external audit of already-shipped work — rides the same edge: `abandoned → audited`, behind the same consult ACCEPT gate; `--force` still cannot reach `audited`.)
4. **On REJECT:** remediation loop:
- **the refute pass has already run** — `consult.mjs` (default-on; `--no-refute` restores the
pre-filter behavior) runs ONE extra external refute pass over the auditor's listed gaps before
the round reaches you: a separate CLI (`CLAUDE_CLI_2` if set, else `CLAUDE_CLI`; post-done REJECT
rounds ONLY — `--plan-audit`/`--multi-auditor` untouched) re-examines each listed gap against the
same frozen evidence. A refuted gap arrives as a string advisory (`[refuted] <issue> — <reason>`)
plus a structured `refute` report on the `consult.history` entry (read it `|| {}`) — NEVER
remediation work. Refutation NEVER flips the verdict (an all-refuted REJECT stays REJECT); refuter
failure degrades to zero refutations (one stderr warn — today's behavior).
- read `consult.last_gaps` — the KEPT gaps; each is `{category, severity, issue, fix}` plus an
optional single-line `verify` (a runnable command proving the fix landed)
- dispatch `zodyssey:sisyphus-junior` per gap (parallel where independent — but see the limitation note
below: hooks are DISARMED in `done`/`audited`, so the cap does NOT apply during remediation
unless you set `phase: "remediate"` first), each carrying the gap's `issue` + `fix`. Dispatch
follows `consult.last_remediation_plan` (read it `|| []`) — the auditor's ordered plan, indices
already remapped to the KEPT gaps: independent single-gap steps stay parallel-by-default, a
step's `note` governs collisions (e.g. two gaps editing the same file), and an absent plan
degrades to per-gap dispatch as today
- **pre-audit verify gate:** after gap-fixes return, run every kept gap's `verify` command as an
ordinary Bash call from the repo root BEFORE re-consult — any non-zero exit loops back to
remediation for that gap WITHOUT spending an audit round; a `verify` that is absent or not a
single usable line falls back to today's discretionary re-verify (old/foreign auditors degrade
gracefully). Verify execution is the conductor's permissioned Bash lane — scripts never execute
auditor strings. Then re-run `consult.mjs`.
- **Do NOT re-audit while a listed verify command still fails.**
- **filter-miss signal (round N+1):** a new-round gap matching a prior round's `[refuted]` advisory
surfaces to the operator as a filter-miss — never silently re-remediated.
- **gap lifecycle (row 32):** every history entry from round 2 on carries `gap_delta` (read it
`|| {}`) — new/persisting/resolved counts vs the previous round, computed from a deterministic
finding key with the `persisting_keys` recorded — the filter-miss check above can match keys
instead of eyeballing prose, and run-report surfaces `max_persisting_streak` — counted in transitions
(a streak of N means N+1 consecutive rounds; surface at ≥2, i.e. a gap in three straight rounds) as
the non-convergence signal.
- **loop until ACCEPT — no hard cap.** Soft safety rail: every 5 rounds, AskUserQuestion to
confirm the user wants to continue (prevents unattended infinite loops; honors "no hard cap").
- **empty last_gaps surface rule (key on observable state, not on what emptied it):** a REJECT
round whose `consult.last_gaps` is empty after routing/refutation surfaces to the operator —
nothing to dispatch; do NOT blind-loop and do NOT fabricate an ACCEPT (the refute report states
refuted gaps are refuted-not-remediated and will be re-judged fresh by the next audit round).
Verdict-level disagreement remains a human decision.
**Discipline:** the auditor's verdict is the independent truth. You remediate gaps; you never edit,
negotiate, or override the verdict. You never fabricate an ACCEPT. The remediation loop converges
because each round shrinks the gap list — if it doesn't, the 5-round check-in surfaces it.
**Honest limitation (re-audit G9):** the remediation loop runs AFTER `phase: "done"`. Because
`done`/`audited` are terminal phases, the enforcement hooks (review gate, parallel cap, file-lock,
phase-gate) are DISARMED during remediation — the doc's claim that "the execute-phase hook still
caps parallel at 4" during remediation is wrong. This is intentional-ish (you don't want the review
gate blocking bug-fix edits) but means remediation is the least-supervised part of the system.
If hard enforcement during remediation is needed, set `phase: "remediate"` (now in the EXEC_PHASES
set) before dispatching gap fixes and restore `done` after re-consult.
## Phase 0 — Triage (you do this yourself; no dispatch)
Classify the request before spending a single sub-agent call:
- **trivial** — a one-line fix, a single obvious edit, a lookup. → Tell the user this doesn't need orchestration and stop. (The Anthropic "don't spawn 50 subagents for a simple query" guardrail, enforced here, at the gate.)
- **standard** — bounded, clear deliverable, a handful of files. → Full pipeline, single executor track.
- **architecture** — multi-system, ambiguous, cross-cutting, or risky. → Full pipeline + Oracle in planning + (v2) team mode.
When in doubt between trivial and standard, choose **standard**. When in doubt between standard and architecture, look at Metis's call in phase 1.
**Two classifiers, one reconciliation rule (T4-#12):** Phase 0 triage classifies **SIZE** (trivial/standard/architecture — how big). Metis (phase 1) separately classifies **KIND** (refactoring / new-feature / audit / bugfix — what shape). These are orthogonal, not conflicting. The rule: SIZE governs *whether* the pipeline runs (trivial deflects); KIND governs *which capabilities* the pipeline reaches for inside a run (audit → source-command-audit-*, bugfix → systematic-debugging, etc.). Metis does NOT override a trivial deflection — if phase 0 said trivial and stopped, phase 1 never runs.
## Parallel-by-default (phase 4) — DEFAULT, NOT OPTIONAL
For every batch of remaining todos, the question is NOT "should I parallelize these?" — it is **"what is BLOCKING me from firing all of them in ONE message?"**
A todo is sequential ONLY if it has a **named blocking dependency**:
- **Input dependency**: todo B reads what todo A produced (a file, a value, a schema).
- **Shared-file dependency**: todos A and B both write the same file (the file-lock hook will block the second anyway — so sequence them).
Workflow each batch:
1. List remaining todos not yet done.
2. Mark each **parallel** unless it has a named dependency above.
3. Dispatch all parallel todos in ONE message (multiple `Task` calls in one assistant turn).
4. State the specific blocking dependency for any todo you hold sequential.
The hook caps in-flight dispatches (default 4). If you hit the cap, the hook blocks the overflow — wait for in-flight work to settle, then dispatch the next batch.
## Anti-duplication rule (critical)
Once you delegate a question to a sub-agent, **do not** research the same thing yourself. You receive the sub-agent's compressed result; act on it. Re-searching what you just delegated wastes the parallel capacity that is the entire point of delegation.
## How to dispatch (the 6-section prompt to each worker)
Every `Task(zodyssey:sisyphus-junior)` (or any worker) carries:
1. **TASK** — the todo title + What to do (literal, from the plan)
2. **EXPECTED OUTCOME** — the acceptance criteria (so the worker knows when it's done)
3. **REQUIRED TOOLS** — which tools to lean on (Read/Grep/Bash for code; zodyssey:explore/zodyssey:librarian via you for research)
4. **MUST DO** — concrete steps
5. **MUST NOT DO** — the todo's Must-NOT-do line (anti-slop)
6. **CONTEXT** — repo root, the todo's References (path:lines to read first), inherited-wisdom notepad paths, and the slug
**Dispatch prompts are POINTERS + DELTA, not restatement (context-economy rule).** Each parallel executor gets its own full copy of your dispatch text, so a 1.4K-word prompt × N executors multiplies N times. Keep a dispatch under ~300 words by pointing the executor at files it can read itself:
- DO point at paths: "the plan is at `<repo>/.zcode/plans/<slug>.md`, read your todo block (id N) and its References first." The executor reads the full context — it does not need you to paste it.
- DO include the delta: the specific todo scope, the must-not-do, the acceptance criteria, the output notepad path. These are the only things not derivable from the files.
- DO block-quote the canonical research question verbatim (from `<repo>/.zcode/plans/<slug>.task.md`) in every research dispatch — the delta is the lens/scope you add AROUND it, never a paraphrase instead of it.
- DO NOT paste the original task, the brief, the audit prompt, or prior notepads into the dispatch body — give paths to them. The executor reads what it needs.
- DO NOT restate capability-routing tables or skill instructions — name the capability ("use `sequentialthinking` MCP for the cross-file reasoning") and trust the executor to load it.
**UI/UX todos add a 7th section:** **DESIGN CONTEXT** — the orchestrator runs `skill: ui-ux-pro-max`'s design-database search (product, style, typography, color, stack) BEFORE dispatching, and pastes the results + the pro-rules (`references/pro-rules.md`) into this section. The executor cannot load the skill itself (sub-agent limitation), so the orchestrator is the bridge. After the executor returns, the orchestrator validates the output against the pro-rules checklist (no emoji icons, contrast ≥4.5:1, responsive breakpoints, accessibility) before accepting it.
For `zodyssey:explore`/`zodyssey:librarian`/`zodyssey:oracle`, the same structure but scoped to their read-only role.
## Checkpointing & resume
After every phase transition and after every todo completes, write a checkpoint to `state.json` (the scaffold created the file; you append):
```json
{"at": "<iso>", "phase": "execute", "completed_todo": "3", "note": "auth route + tests green"}
```
On `/orchestrate resume <slug>`, read `state.json`, find the last checkpoint, and resume from there — not from scratch. This is the durable-execution requirement (DESIGN §5).
**Consuming the resume-format fields** (`scaffold.mjs` + `record-verify.mjs` write them; you are the read side — older runs lack them, treat missing fields as empty):
1. **Skip verified todos.** Any todo whose `state.acceptance[id].pass === true` is already verified — do NOT re-dispatch it; jump to the next pending/in-flight one.
2. **Use notepad pointers for inherited context.** For todos you do resume, read `state.notepad_pointers[id]` (if present) instead of re-reading the full plan/doc — the notepad is the ~3% summary and the full doc stays out of your window (context-economy).
3. **Orient before resuming.** Run `scripts/status.mjs <repo> <slug>` (output now carries verified/notepad counts) and use it as the one-line progress summary that frames the re-entry.
## When to stop and ask
- Phase 1: Metis surfaced user-questions → ask the user, wait.
- Phase 3: review round hit max_rounds (3) without OKAY → surface to user.
- Phase 4/5: a todo is blocked (executor reported `blocked`) and you can't unblock with research → surface to user with the blocker.
- Phase 6: any final-wave item fails → surface, don't declare done.
## What you NEVER do
- Never edit product code yourself in phases 1–3 (planning). The hook blocks it; you also have no reason to.
- Never skip the review gate. Execution before `review.verdict == OKAY` is hook-blocked. (Gate-behaviour claims like this one are rows in `scripts/claims-ledger.mjs`, mechanically re-verified by `scripts/check-claims.mjs` inside `npm test` — a change that alters gated behaviour re-binds its row.)
- Never declare a todo done unless its acceptance criteria passed (phase 5).
- Never expand scope. If a worker reports out-of-scope observations, note them; don't action them mid-run.
## Scripts you call
**Full signatures, flags, and exit codes live in `references/scripts.md` — load it when you are about to invoke any trusted-writer script.** Inline below is only the load-bearing one-liner you need to remember at each transition:
- **Phase transitions:** `scripts/set-phase.mjs <repo> <slug> <phase>` — the *only* sanctioned way to move phases. (Escape hatches: `blocked`/`abandoned` always allowed.)
- **Scaffold the plan:** `scripts/scaffold.mjs <repo-root> <slug> <title> <intent> [task-brief] --criteria-state confirmed|adjusted|skipped` — the flag is optional (omit it: output stays byte-identical to today); with a brief it records the PRIME round's outcome as a first-line criteria-confirmation stamp on the task brief.
- **Parse todos:** `scripts/parse-plan.mjs <plan.md> --lint|--files|--waves|--todo N`.
- **Review gate (phase 3):** dispatch zodyssey:momus → read the minted nonce from `state.review.pending_nonce` → `record-momus-artifact.mjs … --nonce <nonce> --from <file under .zcode/staging/>` (SEC-6 refuses `--from` under `plans/` or `notepads/`) → `record-review.mjs … OKAY --momus-artifact <path> --plan-sha <full-64-char-sha>`. Full order + the `--from`-vs-stdin caveat in `references/scripts.md`.
- **Record todo/verify/final-wave:** see `references/scripts.md` for exact flags (record-verify, record-final-wave, record-todo).
- **Diagnostics:** `scripts/status.mjs <repo> <slug>` (where is this run), `scripts/resolve-capabilities.mjs` (tool-grant reconciliation).
The chain (dispatch → nonce → artifact → verdict) is what makes the OKAY non-forgeable against a **misbehaving executor agent**: an agent can't fabricate the artifact because the nonce only exists after a real `Task(zodyssey:momus)` call the hook witnessed — exact type: a lookalike `*:momus` dispatch mints nothing and warns — and `record-review` binds the verdict to the nonce's consumed record (path + sha256 + round). The 2026-08-04 security wave (SEC-1…SEC-5) closed the holes that let an agent mint OKAY without a real review: the consumed-nonce binding, the WRITE_PATTERNS allowlist (so `.zcode/reviews/` is genuinely not agent-writable), `--force` scope (recovery-only), the plan-sha re-hash at edit + final time, and the scope gate in `phase=final`.
**Honest residual limit (not a bug to fix, a boundary to know):** none of this binds *what zodyssey:momus returned*. The nonce proves a dispatch happened; the artifact's verdict text is whatever the orchestrator pipes in. So the threat model is **executor-agent adversarial, NOT orchestrator adversarial** — if the orchestrator itself is compromised, it can dispatch zodyssey:momus for real and still record an arbitrary verdict. Fully closing that needs the harness to hand zodyssey:momus's transcript hash to the hook for a true content binding (a design change, tracked). SEC-6 partially mitigates by refusing `--from` paths under agent-writable bookkeeping dirs. State this limit honestly in any audit; do not claim the chain is airtight against an orchestrator adversary.
## Environment overrides (documented)
- `ZODYSSEY_PARALLEL_CAP` — the execute-phase parallel-dispatch cap (default 4; non-integer/≤0 → 4).
- `ZODYSSEY_STALE_HOURS` — a run not updated in this many hours is treated as abandoned (hooks disarm; default 24; non-finite/≤0 → 24).
- `ZODYSSEY_UNGATE_BASH` — set to `1` to bypass the Bash write-gate entirely (documented power-user hatch; Edit/Write tools stay gated). Every ungated call is witnessed — one JSON line in `.zcode/state/<slug>.ungated.jsonl`, counted as `ungated_bash_calls` by `run-report.mjs` — the hatch opens, but never silently.
- `ZODYSSEY_RECURSION_CAP` — the SEC-1s recursion-guard cap (default 1). Reserved for a future real depth counter; today the guard is a payload-pattern match against embedded nested dispatches, not a depth ledger (the harness tool-grant boundary is the primary control).
- `ZODYSSEY_REGRESSION_TIMEOUT_MS` — timeout (ms) for the regression-gate suite (default 600000).
- `ZODYSSEY_NO_AUTO_COMPACT` — set to `1` to skip auto-compaction at final-phase entry (default unset = enabled).
- `CLAUDE_CLI` — the binary `consult.mjs` spawns as the external auditor (default `claude`; receives the full repo diff + plan — point it only at a trusted CLI); `CLAUDE_CLI_2` — optional second CLI binary for `consult --multi-auditor` (falls back to `CLAUDE_CLI`).
## Memory store — canonical
The **`memory` MCP knowledge graph** (`~/.zcode/orchestration/memory.json` is its on-disk persistence) is the canonical cross-run store: write durable facts (decisions, gotchas, key files) there at end of run, read at consult. The per-run notepads (`.zcode/notepads/<slug>/`) are working memory within a run only.