Use at session start (and after compaction) to resume an interrupted atlas-orchestrator run — whether a single-change `atlas` run or a multi-node `atlas-weave` graph run. If the cwd holds an unfinished `.atlas/<run_id>/` ledger, pick up from the durable on-disk state instead of restarting. Safe no-op when there is no `.atlas/` here.
Installs into .claude/skills of the current project.
Are you the author of Atlas Resume?
Add the live security badge to your README. It updates with every re-scan.
[](https://www.skillsdirectory.com/skills/null0xxx-atlas-resume-atlas-orchestrator)
---
name: atlas-resume
description: Use at session start (and after compaction) to resume an interrupted atlas-orchestrator run — whether a single-change `atlas` run or a multi-node `atlas-weave` graph run. If the cwd holds an unfinished `.atlas/<run_id>/` ledger, pick up from the durable on-disk state instead of restarting. Safe no-op when there is no `.atlas/` here.
---
# atlas-resume — on-disk run resumption (F1 / P11)
**Invoke this skill directly whenever you suspect an unfinished run — this manual invocation is
the MANDATORY, load-bearing recovery path.** In the opencode port there is **no SessionStart hook
and no automatic pointer** (the Claude Code port's `hooks/session-resume.sh` does not exist here),
so this manual invocation is the *only* recovery path — never merely a backup to an
assumed-working hook. Do not wait for any automatic signal to appear before resuming by hand; run
this skill on suspicion alone.
This is a **pure instruction**. It injects **no live state** — it only tells you *where the durable
ledger lives on disk* so you can find it yourself. atlas-orchestrator keeps its authoritative run state on
disk (never in context), because the full orchestrator prompt is **not guaranteed to survive
compaction**. This skill body IS re-injected at session start and after compaction, so it is the
reliable pointer back to that on-disk state. On smaller-context models a multi-node run compacts often (the
root reads every node's return into its own context), but **on large-context models compaction is
RARE** — this resumption then covers turn-kills and crashes more than compaction. It stays
load-bearing (correctness must survive it either way), just no longer the common path on those models.
## What to do at session start
**Script-call convention** — every script call in this file (both paths below) runs as
`python3 -c "from scripts import <mod>; …"`. The **INIT export** convention (see the `atlas`
skill's platform-facts section) sets `PYTHONPATH` (**set to** `${ATLAS_PLUGIN_ROOT}`, the plugin
root, **and nothing else**, so `from scripts import <mod>` resolves against the plugin rather than
against any directory the environment names), `PYTHONSAFEPATH=1` and `PYTHONNOUSERSITE=1`,
exported ONCE at the start of a run into opencode's **persistent `Bash` shell session**. If this
skill runs standalone at session start (no atlas run has exported them yet), run that export
FIRST — before any script call below — then locate the newest ledger. `PYTHONSAFEPATH=1` remains
mandatory in that exported environment: the working directory
here is the user's own project, and without it that directory would outrank `PYTHONPATH`, letting
the project replace `scripts/ctxstore.py` or `scripts/resume.py` with its own. `PYTHONNOUSERSITE=1`
closes the channel the other two do not reach: `site` imports `usercustomize` from the user site
directory **at startup**, so a `usercustomize.py` planted through an ambient `$PYTHONUSERBASE` runs
inside the interpreter before any of this skill's own code does. **The claim stops there, narrower
than it used to read:** `$PYTHONHOME` relocates the stdlib itself and is **still open
session-wide**, so "cannot be steered by whatever the environment held" would be false — three of
the four resolution channels are closed, not all four. Do **not** add a
per-invocation prefix back: a reintroduced prefix would shadow the session's correct values with a
broken relative path instead of reinforcing them.
1. **Look in the current working directory only.** If there is **no `.atlas/` here, do nothing** —
stop silently and proceed with the session normally.
2. **Decide graph-run vs single-change.** A **graph run** is one whose `.atlas/<run_id>/` holds a
`plan.dag.json` (the `atlas-weave` outer machine). A **single-change run** has only the
`atlas` ledger (`state.json`, no DAG). Discover the runs on disk (each `.atlas/*/` with its
`state.json` `current_state` + mtime + whether a `plan.dag.json` exists), then:
### Graph run (atlas-weave) — re-derive the frontier by pure projection
3g. **Select the graph ROOT run.** Call `resume.select_graph_run(runs, session_id)` with the on-disk
run descriptors (`{run_id, has_dag, state, mtime}`). It returns the non-terminal run that carries a
DAG and is **not** a task sub-run (`resume.is_task_subrun` skips any `${SESSION}/tasks/<id>`
sub-run), preferring the current session, else the newest by `(mtime, run_id)`. If it returns
`None`, there is no resumable graph run — fall through to the single-change path or stop.
4g. **Reset the orphaned frontier.** Read `plan.dag.json`; `dag = resume.resume(dag)` — this resets
every orphaned `RUNNING` job (its inner-atlas agent died with the turn) back to `PENDING` and
clears its lease, WITHOUT `attempts++` and WITHOUT refunding gas (a compaction is not an agent
failure, and the interrupted dispatch already spent its fuel — so re-dispatch stays gas-bounded
and the run still provably halts). Terminal (`DONE`/`FAILED`) and `PENDING` jobs are untouched;
no node is dropped. Write the reset DAG back with `ctxstore.write_artifact_atomic`.
5g. **Discard in-flight receipts (the lease no-rotation rule).** The lease token
`f"{job_id}#{attempts}"` does NOT rotate across this reset (attempts is unchanged), so a receipt
from the killed turn would still pass `scheduler.lease_valid` against the re-dispatched attempt —
**ignore any such receipt.** Only receipts produced *after* this resume count.
6g. **Reset dirty worktrees.** Any per-node or union worktree left half-written by the killed turn is
untrusted — remove it (`uniontree.cleanup` / `git worktree remove --force`); the re-dispatched
node will re-create a clean one at its baseline.
7g. **Re-enter the outer machine.** Resume the **atlas-weave** machine at **SCHEDULE** (the frontier is now
re-derived) with the **same** `run_id` and the same frozen packet + `success_criteria` (never
re-derive them). Continue draining the pool → INTEGRATE → AGGREGATE → OUTPUT. Honor every gate:
never auto-apply the union; stop at the OUTPUT gate exactly as a fresh run would.
### Single-change run (atlas) — the original path
3s. **Find the newest unfinished run** among `.atlas/*/state.json` with `current_state != "OUTPUT"`.
If every run is at `OUTPUT` (or none exists), **do nothing** and proceed normally.
4s. **Read the ledger, do not restart.** Read the immutable `intent` + frozen `success_criteria`
(never re-derive), the `stages` ledger + `current_state`, `refine_passes`, `verify_cmd`,
`scope_paths`, `baseline_sha`, and `log.jsonl`.
Before 5s, inspect the selected run's structured `clarify_resolution`. If TRIAGED is not done and `triage_ready(load(raw))` is false, resume CLARIFY as described below instead of taking the successor of its last ledger entry. CLARIFY is persisted before waiting; its `done` ledger marker records durable interview progress, not user confirmation. Do not skip pending frontier questions or confirmation because this marker exists.
5s. **Resume from the last recorded stage.** Re-enter the **atlas** state machine at the stage
**after** the last one recorded `done`, in the **same** run (same `run_id`) — **with the one
exception below**. Do not start a new
run, re-run completed stages, or re-capture intent.
> **`REFINE` is the one entry whose successor is NOT the next `STAGES` member.** A last recorded
> `REFINE` means the gate already decided `REFINE?=True`, so resume by **re-entering the refine
> loop at `CODED`, never `OUTPUT`** — re-dispatch the coder, then `VERIFIED`. Resuming at
> `OUTPUT` would print a status computed from the verdict that decision had already superseded:
> the forced pass never ran, and the run reports ✅ for a tree nothing re-verified. Reaching
> `OUTPUT` directly from a trailing `REFINE` is legal **only** on the degraded could-not-verify
> path (the coder could not be re-run at all), and that path requires `budget_exhausted = True`
> at OUTPUT — i.e. ⚠️ UNVERIFIED, never a green. `floorsynth.stale_verdict_defects` blocks the
> shape either way, so a resume that skips `CODED` cannot be laundered into a ✅.
>
> *(A2: copied verbatim from the Claude Code orchestrator skill, which has carried this
> prohibition all along. This file did not — and a resume skill is read precisely when context
> was lost. `REFINE → CODED` is already legal in `fsm.py`; no new `advance()` is involved.)* The pass counter is the count of `REFINE`
entries in the ledger, read from disk, never from memory.
6s. **Honor the run's gates.** Never auto-apply to a real tree; stop at the pre-CODE approval gate and
the OUTPUT gate exactly as the `atlas` orchestrator would.
## Safety
A malformed structured interview in an otherwise selected run requires recovery; it is not permission to abandon its unanswered decisions or treat it as confirmed. For other cases, if a ledger/DAG is unreadable, treat this as "no resumable run" and proceed
normally — resumption is best-effort and must never block or corrupt a fresh session.
## Resume an interrupted CLARIFY interview
For the selected single-change run, read `clarify_resolution` before choosing the next stage. A structured interview resumes through `scripts.grilling_state.load(raw)`: restore its tree, settled answers, recorded rounds, and stored frontier. Continue the same full frontier with recommendations; pending exploration blocks only dependent nodes. After answers use `record_round(state, frontier_ids, answer_strings)`, reshape open branches via `set_tree`, fold settled answers into still-mutable packet fields and revalidate them, and persist the actual `dump(state)` through `ctxstore.advance(..., "CLARIFY", updates={"clarify_resolution": grilling_state.dump(state)})`. Never insert user text into Python source or replace this data with a placeholder.
If all branches are settled but `confirmed` is false, obtain explicit human confirmation of shared understanding. Only the actual reply permits `confirm(state, user_confirmed=True)`; persist it before TRIAGED. `triage_ready(state)` and the shared ctxstore gate prevent freezing an unfinished interview. A headless run remains UNVERIFIED. Malformed or unknown structured records require recovery; never silently reinterpret them as an answered interview. Historical plaintext resolutions retain their previous path.