Drive Splotch's deployment-target performance matrix from current evidence to zero unexplained scoreable red cells on the release-gate rows through product improvements and faithful recaptures, keeping harness work subordinate to and immediately useful for a named product experiment. Ships each causal product cluster as its own reviewed PR, merged before the next cluster begins; an unattended run goes through ship-campaign profile=performance. Use for sustained performance improvement; use ca...
Installs into .claude/skills of the current project.
Are you the author of Improve Performance Matrix?
Add the live security badge to your README. It updates with every re-scan.
[](https://www.skillsdirectory.com/skills/kylemit-improve-performance-matrix)
---
name: improve-performance-matrix
description: Drive Splotch's deployment-target performance matrix from current evidence to zero unexplained scoreable red cells on the release-gate rows through product improvements and faithful recaptures, keeping harness work subordinate to and immediately useful for a named product experiment. Ships each causal product cluster as its own reviewed PR, merged before the next cluster begins; an unattended run goes through ship-campaign profile=performance. Use for sustained performance improvement; use capture-performance-matrix for capture-only snapshots or validation.
---
# Improve performance matrix
Run a fresh evidence-led campaign against the authoritative deployment-target matrix. The campaign
ends only when every scoreable cell on a **release-gate row**, however old its capture, is green or
carries a recorded, evidence-backed disposition, unless the user sends a control message that
explicitly requests a stopping point (**wrap up** or **stop at mergeable**, below). ADR-0156 defines
the rows: the physical iPad (web and native) and the physical Android phone (web and native) are the
release gate; Mac rows are a regression tripwire; simulator and emulator rows are advisory and never
count toward completion. ADR-0160 defines the disposition: an ADR-recorded measured allowance or
documented floor that names the cell's measured basis, its trace attribution, and the condition that
reopens it. A red cell with one is **explained** and counts toward completion; the campaign's
remainder is the **unexplained** reds. A disposition is a release-gate policy change the owner
records, never something a campaign grants itself to finish.
This is the improvement sibling of `capture-performance-matrix`: that skill owns comparable capture
mechanics and matrix refreshes; this skill owns inventory, causal attribution, product optimization,
capture-path repair or faithful recapture, merge-as-you-go delivery, review, and campaign control.
For an unattended run, `ship-campaign profile=performance` supplies the preflight, the ledger, the
quarantine rules, and the morning report around the cluster loop this skill defines.
An explicit user request to run this improvement campaign authorizes its normal in-repository
branches, commits, pushes, PRs, rival reviews, `start-capture-session` device reservation, and
merging each cluster's PR through `ship-issue`'s autonomous merge gate. Merely loading the skill for
planning or reference authorizes none of it. Before publishing the ledger, launching an external
reviewer, or starting unattended captures, complete `ship-campaign` step 1's consent preflight.
Reuse the user's existing explicit approvals and ask one bundled question for missing scope:
publishing code, device context, and measurements to the named repository host; sending repository
code/diffs, device context, and measurements to the rival's provider for consultation and PR reviews
(Anthropic Claude from Codex, OpenAI from Claude); merging reviewed, passing PRs; and the unattended
stop time and timezone or explicit continuous-goal completion condition. Quote the user's actual
words in the ledger and each cluster's authorization block. Verify publishing and reviewer approval
through their first real authorized operations, recording acceptance or the classifier's stated
rejection; installation health alone does not verify that consent. No quote pre-approves every
future operation, and no overnight grant authorizes extending the recorded stop time indefinitely.
## The campaign advances one merged PR at a time
The scheduling loop is:
```text
bounded product pass → verify → commit → open PR → rival round 1 → address → rival round 2 → address → green CI → merge → next pass from fresh main
```
Opening, reviewing, and merging the PR are part of completing the product pass, not end-of-campaign
shipping. Do not begin another accepted treatment or accumulate more product commits while either
required rival round, CI, or the merge is outstanding. Every pass branches from the `main` that
already contains the previous cluster, so an early finding is fixed in the PR that introduced it and
a mistaken premise cannot compound through later experiments. A stack of unmerged clusters is used
only when the user asks for one (`create-stacked-prs`).
## Product work is the deliverable
This is not a harness-improvement campaign. Splotch's profiling harness is mature and presumed
sufficient. After one current, scoreable failure is reproduced on a calibrated physical release
gate, the next meaningful output is a product hypothesis and A/B experiment. Harness work may be
part of that experiment when it directly improves diagnosis, execution, or validation; broader
capture coverage, richer metadata, another metric, or generalized tooling is not an alternative to
the product loop.
Read the whole published matrix at the start, but do not block the first product experiment on
recapturing every old section. Coverage of the release-gate rows remains part of the completion gate
(ADR-0156), and so does their age: none may be older than `RELEASE_GATE_MAX_AGE_DAYS` when the
campaign finishes (ADR-0175), so an old release-gate section waits for completion, not for the first
experiment; advisory rows are recaptured for breadth when the rig is free, never as a completion
requirement. Coverage is not a prerequisite for beginning product work when a calibrated physical
target already provides a reproducible failure. Old advisory Simulator, emulator, and desktop rows
cannot delay that first experiment. Recapture an old authoritative section first only when its
result is necessary to distinguish the selected hypothesis or establish a calibrated failure.
Treat `tools/perf/`, matrix schemas and generators, capture transports, scorers, evidence formats,
and profiling documentation as stable supporting infrastructure while working a product cluster. A
harness change is allowed when all of these are true:
1. The live campaign ledger already names the current product failure, the product hypothesis, and
the before/after experiment the harness change will serve.
2. The existing path was attempted or inspected closely enough to show a concrete shortcoming for
that experiment. The change may repair invalid or incomparable evidence, add a diagnostic needed
to distinguish the hypothesis, or make the faithful A/B materially more reliable. Convenience,
polish, future reuse, or an isolated advisory-runtime anomaly is not enough.
3. The proposed change is the smallest useful slice, and the same cluster uses it immediately;
generalized cleanup and adjacent improvements become issues rather than campaign commits.
4. The campaign immediately returns to the same product experiment, records its product outcome, and
does not promote the harness change itself as a delivered causal cluster.
If harness work starts expanding beyond what the named experiment will use immediately, stop and ask
the user before expanding scope. A checkpoint with current product reds and no product outcome is an
**inventory checkpoint; campaign incomplete** — never a completed improvement, shipped cluster, or
merge-ready campaign result.
Keep the PR ledger honest about allocation. At every status update, list product commits, harness
repair commits, and capture/evidence-only commits separately, name the product experiment each
harness commit served, and say explicitly when no product optimization has landed.
## Start from current truth
Do not inherit a red-cell count, causal theory, target list, or optimization priority from an older
campaign prompt, PR body, report, or memory.
1. Preserve unrelated local work. Do not stash, delete, clean, commit, or absorb it. If the working
tree is not clean, stop and tell the user rather than carrying their changes onto a campaign
branch.
2. Inspect open performance PRs before creating anything. If an unfinished campaign PR already owns
the matrix work, verify its branch, PR, and checkpoint state and finish it instead of duplicating
it. Otherwise fetch the trunk, verify prior campaign PRs are merged, and branch the first cluster
from a fresh `origin/main`.
3. One accepted cluster is one PR. Open it as a draft as soon as its first coherent commit is
pushed, complete both rival rounds, address their findings, get its CI green, and merge it before
starting the next cluster from the updated `main`. Never accumulate multiple accepted clusters on
one branch for later decomposition; rejected or inconclusive experiments stay local and are
backed out. The campaign invocation already authorizes these PRs and merges, so do not wait for a
later request to create them.
4. Keep the live campaign ledger as a single comment on the performance tracking issue, edited in
place (the `ship-campaign` ledger; a scratch file only when no tracking issue exists). Record the
baseline inventory, merged clusters, current cluster, remaining work, exact product commits, raw
artifact provenance, correctness evidence, and matrix status. Each cluster's PR body carries that
cluster's own evidence.
5. Read the `profiling`, `capture-performance-matrix`, and `testing` skills. Read `mobile` before
any iOS, Android, or Capacitor work.
6. Locate the authoritative matrix inputs, source manifest, and generator from the current
repository rather than carrying paths or output names forward from an older campaign. Discover
generator-owned JSON, Markdown, and HTML outputs from the generator or directory instructions.
`scrapbook/performance/` contains several matrices: identify the deployment-target matrix by its
scope, source manifest, and generator rather than assuming the newest or first `data.json` is
authoritative.
Resolve the authoritative deployment-target matrix first, then inventory every published cell in its
`data.json` before editing. This is a read-only classification pass, not a requirement to recapture
every old cell before product work. Classify each cell by:
* target and deployment class;
* web or native runtime;
* orientation and theme;
* drawing brush, undo case, or discrete action;
* failed metric and raw values;
* scoreability and control validity;
* capture age and exact product commit;
* capture source, runner, input transport, and raw provenance;
* freshly captured versus preserved section;
* comparable versus historically invalid instrumentation.
Report these categories separately:
* genuine product failures, each with its capture age;
* old captures a hypothesis needs refreshed before it can be judged;
* preserved historical captures;
* invalid or unscoreable modes;
* incomparable captures;
* runner or capture-path blockers.
Generate counts from live data, not prose. Before treating any red cell as a product problem, run
`npm run check:matrix-staleness -- --base=origin/main`. It ranks every section by capture age and
counts the engine and product commits that landed since (ADR-0175). A red cell describes the commit
it was captured at, and the product moves underneath it: the 2026-08 campaign wrote five candidate
implementations against a gate a prior extraction had already fixed. The check needs no device and
answers in seconds. The explicit `--base` matters: the default is `HEAD`, which from a campaign
branch counts the branch's own commits as drift. An old red still counts toward the gate, but when
product commits on its measured path landed since, recapture it before building a product hypothesis
on it.
The same check applies to numbers a campaign **prompt** calls established. A prompt is written from
the matrix and the sessions before it, so its "measured" figures carry the commit they were measured
at, not the trunk's; on 2026-09-02 a prompt's central cause (an ~86 ms `clear coloring page` raster
on every physical iPad cell) had been fixed on `main` the day before by a commit the prompt's author
never saw, and a full layer was built and A/B-tested against it before a concurrent control on
`main` showed the cell already green. Run the age report and one concurrent control on the trunk
before building on a prompt's figures, however authoritative their framing.
As soon as a current calibrated physical failure exists, turn the remaining genuine failures into a
compact causal-cluster inventory with a representative cell, affected blast radius, evidence
confidence, and next discriminating product experiment. Select one and start its product A/B.
Prioritize a systemic cause that plausibly explains several cells over isolated tail-chasing, but
let current evidence choose the order. Continue recapturing old sections for breadth and completion;
do not use it as a blanket reason to postpone the selected experiment.
## Non-negotiable evidence rules
Never make the matrix green by:
* relabeling a failure, weakening a gate, or changing scoreability to exclude it;
* reducing visual, input, drawing, native, or export fidelity;
* skipping actions, brushes, themes, orientations, runtimes, or difficult samples;
* publishing only a lucky retry or discarding a faithful red result;
* copying a pass from another target or calibration tier;
* treating old, incomparable, or invalid evidence as approval of the current product.
Frame pacing and readiness are separate acceptance dimensions. The action scorer's `passed` verdict
covers first response and presented-frame continuity; `readyMs` records when the action-specific
observable outcome actually arrived. A cluster is not accepted from a greener frame verdict alone.
For every discrete-action A/B:
* compare readiness P50/P95 from the same action, target, runtime, transport, polling cadence, and
ready predicate, alongside first-frame and post-action distributions;
* reject a candidate that moves required work beyond the scored activity window, weakens the ready
predicate, or delays observable completion merely to protect animation frames;
* treat a readiness regression larger than the capture path's measured resolution/noise as a product
tradeoff, not a performance win. Keep it only with explicit user approval and record the frame
benefit, latency cost, and why the deferred work is non-critical;
* when an intentionally deferred action remains, apply an activation/busy state synchronously and
keep it visible until completion. A non-idempotent activation must be single-flight; repeatable,
idempotent choices such as selecting the current color need no artificial input lock;
* capture normal-speed before/after video or GIF for any changed temporal behavior, cropped to the
control and affected surface, before asking for the appearance verdict.
The committed matrix reports readiness P95 but does not assign one universal gate: “ready” ranges
from a local state flip to a full-resolution download, and remote drivers add different polling
floors. That is why the comparable A/B requirement above is mandatory rather than an invitation to
ignore the number. If a capture path cannot resolve the proposed readiness difference, it cannot
approve that experiment; use a finer in-page mark or another faithful path immediately serving the
named product hypothesis.
An old red cell clears only through a faithful fresh capture or a recorded disposition. Harness work
follows the product-first gate above: repair a demonstrated measurement defect or add a targeted
diagnostic or validation capability only when the named product experiment will use it immediately.
Do not create a freestanding harness roadmap inside the campaign. Promote representative raw
captures with
`npm run perf:evidence:keep -- --corpus=<dir> --campaign=<name> --product-commit=<capture-product-sha>`
(the campaign directory or one `<campaign>/<target-id>` directory; add `--target=<id>` only when no
path names the target) so they remain rescoreable, and trial a scorer change across that preserved
corpus with `npm run perf:rescore -- --corpus=perf-profiles/evidence/<name>` before treating it as
valid. Preserve or strengthen coverage and add a regression test for the exact measurement failure.
Preserve drawing output, undo semantics, coloring selection and clearing, settings and persistence,
rotation restoration, export fidelity, native/web parity, accessibility, and toddler-facing visual
and interaction feedback. A faster incorrect interaction is a rejected experiment.
That rule has a scheduling consequence: **when a candidate change alters what the user sees — brush
texture, deposition, color, animation — get the human appearance judgement before spending device
time on its timing campaign.** Two 2026-08 candidates produced fidelity-passing, CI-green,
independently reproduced timing wins that were then declined on appearance ("reads as a glitch
rather than as ink drying… whatever the frame numbers say", ADR-0147); the wasted captures were the
only waste class in that campaign where every measurement was correct. A screenshot or short
recording for the user costs minutes and no device occupancy; ask for the verdict early and run the
timing campaign on candidates that already look right.
## Physical-device boundary
Before the first physical capture, use `start-capture-session`, read `docs/PROFILING-CAMPAIGNS.md`
completely, and run:
```sh
npm run perf:preflight -- --wake-android --verify-android-input --verify-ios-launch
```
Then follow these invariants:
* discover and prove device endpoints and ports dynamically;
* serialize all physical-device captures and keep the host otherwise quiet;
* never kill, attach to, or reuse a foreign process merely because it owns a preferred port;
* build fresh instrumented artifacts from the exact product commit before measuring;
* bind every folded capture to its product commit, built entry, runner, transport, target, and raw
source;
* never commit device identifiers, local capture scaffolding, credentials, transient reports, or
machine-specific state;
* ask once with the exact action only when unlocking, Guided Access, XCTest authorization, or
another genuinely human-only device interaction blocks progress;
* capture one known-good control cell after the preflight goes green, before the first experimental
capture — a cell whose expected value is on record. A control that lands in band proves the whole
path (build, serve, transport, fidelity, scoring) in one capture; a surprising first experimental
number on an unproven path is undiagnosable.
Simulator, emulator, desktop, and uncalibrated results are advisory unless the current profiling
rules explicitly give them approval authority (ADR-0156 names the release-gate rows). Use them to
reject or narrow hypotheses; do not let their passes overrule a calibrated physical failure.
On both physical iPad rows, drawing lost-frame share is judged against the real-finger floor, not
the driven capture (ADR-0174). XCUITest touch synthesis adds about one point in Safari. A driven pen
reading inside ADR-0174's recorded band (above 1%, up to 1.37%) is already explained, so do not
spend product work on it. Any other new driven iPad drawing lost-frame red, including a Magic
reading at any commit or an eraser reading at a commit ADR-0174 does not name, needs a
`perf:device:hand` capture at that commit before you count it as a product red or an artifact. The
eraser reds at e5142fab and 3928cd88 are already explained by finger captures, so do not reopen
them. The Magic first-load stall is recorded as not reproduced at 8e6700d5; it reopens only if a
real-finger Magic capture shows an in-contact frame over 50 ms.
## Work one causal cluster at a time
For each cluster:
1. Reproduce the smallest representative failure on the current product and capture path.
2. Inspect raw traces, activities, input delivery, engine marks, frame intervals, layout, paint,
raster, GPU/compositor work, and capture metadata — not only the final pass boolean.
3. Attribute the expensive frame or invalid result to a concrete product, runner, transport, or
instrumentation cause. Separate first-action latency from post-action work and physical
corroboration from simulator-only behavior.
4. If attribution remains ambiguous, first use existing narrowly scoped diagnostics or a supported
user-flow A/B control. Any harness edit must pass the product-first direct-utility gate and must
immediately return to this same cluster. Do not change the scorer to hide ambiguity.
5. Make one causally coherent product change. Back out rejected or inconclusive experiments instead
of stacking speculation; a measured rejection is still the product outcome that closes the
experiment loop.
6. Run focused correctness tests and the exact failing performance case, plus `npm run check`,
`npm run lint`, and `npm run format:check` before any commit that touches code or scripts.
7. Compare raw before/after traces and generated summaries from faithful runs. Preserve the first
valid red after a change rather than retrying it away.
8. Verify the real app visually and behaviorally, including every interaction contract the change
can affect.
9. Broaden across all affected themes, orientations, brushes/actions, and web/native targets.
10. Recapture complete affected modes, not only the original sample.
11. Fold only faithful, comparable captures into the authoritative matrix. Mark every captured row
the campaign did not recapture `preserved`, then regenerate with
`npm run gen:performance-matrix -- --strict <manifest>`. That command runs the age report
in-process against the manifest it resolved, and `--strict` fails any section that lacks a
`capturedOn` date or a resolvable product commit (ADR-0175). Validate every generator-owned
output and prove JSON/Markdown/HTML agreement where present.
12. Commit and push each causally coherent verified product improvement separately, update raw
evidence and remaining status in the campaign ledger, and proceed only from a clean tree. A
directly useful harness change may precede it in the same cluster, but never substitutes for the
product outcome it exists to support.
## Review and merge discipline
Deliver each causally distinct product cluster as its own PR from a fresh `main`, and put it through
`drive-pr-to-mergeable`. On a shippable verdict, merge it under the all-or-nothing gate `ship-issue`
step 5 defines — live re-verification of the head, a rival review that actually posted, every
applicable check green, the merge commit copied from command output — and confirm the merge on
`origin/main` before the next cluster begins. A not-shippable verdict stops new clusters and reports
the blocker; under `ship-campaign` it quarantines the cluster instead. `drive-pr-to-mergeable` owns
the reviewer, the CI loop, and the verdict; this campaign adds only the following.
Do not create a standalone harness-improvement cluster unless the user explicitly asks for one; an
incidental repair stays subordinate to the product cluster it unblocks. Every PR body includes:
* root cause and causal scope;
* exact raw before/after metrics and artifact provenance;
* exact product commit used for each capture;
* capture target, runtime, runner, transport, and fidelity verdict;
* correctness, visual, parity, persistence, rotation, and export checks that apply;
* current matrix status and explicitly remaining clusters.
**Reviewer budget override.** Invoke `drive-pr-to-mergeable` with round two **unconditional**: run
it even when round one found nothing, resumed so the rival can verify its own disposition —
performance changes can preserve behavior and still encode a mistaken causal theory. A material fix
landing after round two reports — whether prompted by that round, a human comment, or CI — earns one
more resumed verification round inside the rival's three-round budget; if that round finds another
material issue, address it and use `--fresh` for one final pass; if the fresh pass finds another,
stop starting clusters and report the repeated-review blocker rather than extending the loop.
A substituted, same-runner reviewer withdraws the merge authority, as in `ship-issue`: the cluster
finishes as an open, mergeable PR for the user, and the next cluster waits for it rather than
stacking on it.
## Control messages
Treat these as steering inside the active campaign, not as replacements for the campaign objective.
* **status** — report the overall campaign status in commentary, including the freshly established
baseline; product clusters and PRs already merged; product, harness-repair, and
capture/evidence-only commits as separate lists; the product experiment each harness repair
served; the exact current in-flight cluster, phase, branch/PR, and latest evidence; remaining open
red cells with their capture ages, and incomparable, unavailable, and blocked cells; current
runner/device blockers; and the best evidence-based estimate of work left. State explicitly when
no product optimization has landed. Re-read live matrix, git, PR, review, and CI state where it
may have changed. Do not send a final response, stop tools, pause captures, or treat the question
as turn-terminating. Continue the in-flight campaign after answering.
* **pause** — stop selecting new clusters, finish or safely back out the current experiment, leave
the current branch and PR evidence coherent, push a recoverable checkpoint, and report the exact
resume point. Do not merge or present a paused partial cluster as mergeable.
* **resume / continue** — verify live matrix, branch, PR, artifact, device, and CI state before
resuming the recorded cluster. Do not assume the previous process, port, build, or capture remains
valid.
* **wrap up** — stop selecting new clusters, but completely finish the current in-flight cluster as
shippable work: resolve attribution, land or back out the experiment, run focused correctness and
exact performance validation, complete affected-mode recapture when needed, fold only faithful
evidence, regenerate authoritative outputs, commit and push, update the live campaign ledger,
complete review and feedback, drive CI green, and merge it through the gate above — or back it out
and leave the PR as a draft with its evidence when it cannot pass. Then report both the merged
scope and the freshly counted campaign remainder; do not claim the overall matrix is complete when
cells remain. Wrap up never resumes the campaign.
* **stop at mergeable** — everything wrap up does except the merge: drive the in-flight cluster's PR
to a shippable verdict and leave it open for the user, reporting it as the next merge. Use it
whenever the user asks to make work mergeable, ready, or reviewable without asking to land it —
merging is irreversible, so it is never inferred from a request to prepare.
A casual progress question such as “what is running?”, “where are we?”, or “how much is left?” is a
**status** message. Phrases such as “finish what is in flight” or “stop after the next complete PR”
are **wrap up** messages, and “make this mergeable” or “get it ready for me” are **stop at
mergeable** messages, unless the user explicitly asks to continue to zero.
## Optional Goal mode
The workflow must not depend on provider-specific goal tracking. The matrix, raw artifacts, git
history, merged PRs, and campaign ledger remain the durable source of truth.
When Goal mode is available, use it only if the user explicitly requests Goal mode for this
campaign. Create one objective for zero scoreable, unexplained red cells on the release-gate rows,
each shown with its capture age, and no release-gate section older than `RELEASE_GATE_MAX_AGE_DAYS`
(ADR-0156, ADR-0160, ADR-0175), and omit a token budget unless the user supplies one. Goal mode is
useful for automatic continuation and for keeping the terminal condition visible across long tool
runs. It is a poor fit for an ordinary campaign that may receive `pause` or `wrap up`: it supports
completion or genuine blocking, not a wrap-up or stop-at-mergeable pause, permits only one active
goal, and does not replace external checkpoints. Never mark the goal complete for an improvement, a
green cluster, or a wrap-up that leaves scoreable, unexplained reds on a release-gate row, however
old.
## Completion gate
Complete the full campaign only when:
* a freshly regenerated matrix has zero scoreable, **unexplained** red cells on the release-gate
rows, each shown with its capture age (ADR-0156, ADR-0175) — the matrix's **Open release-gate
reds** list is that count, and an old red keeps counting until it is recaptured or explained. A
red cell counts as explained only when an ADR records its disposition with the measured basis, the
trace attribution, and the reopen condition (ADR-0160's measured allowances are the shape; the
matrix renders every allowance beside its gates), and a disposition granted by the campaign itself
rather than recorded by the owner does not count — and no release-gate cell that is unscoreable
because its instrument is uncalibrated — such a cell counts as red until the runtime is calibrated
or recorded as uncalibratable; simulator and emulator red is rendered and reported, never counted
as remainder, and a Mac cell counts only when it turned red on a change that was green on the
trunk;
* every genuine product red on a release-gate row (or a Mac cell that turned red on a change that
was green on the trunk) that existed during the campaign has a recorded product outcome — a
verified improvement or an empirically rejected candidate followed by the next hypothesis; a
campaign with such reds and only harness, documentation, or capture commits is incomplete;
* every unavailable scoreable cell on a release-gate row has a faithful capture, and every captured
section has a `capturedOn` date and a resolvable product commit (`--strict`, ADR-0175); an
advisory row the campaign did not recapture is marked preserved;
* every release-gate section is at most `RELEASE_GATE_MAX_AGE_DAYS` old
(`tools/perf/lib/capture-date.mjs`): `npm run check:matrix-staleness -- --release-gate-age` exits
0 (ADR-0175, as amended). Recapture each section it names; a product fix brings a fresh capture
for free, so this catches only the sections no cluster touched. Tripwire and advisory rows have no
age limit;
* capture-path blockers on release-gate rows are fixed and every affected release-gate target is
recaptured; an advisory target's blocker is filed as an issue, and a genuinely unsupported mode
stays explicitly unscoreable rather than being counted as a pass;
* correctness, accessibility, visual behavior, native/web parity, persistence, rotation, undo, and
export fidelity remain intact;
* every authoritative generated output agrees;
* every campaign PR is reviewed and merged with its merge commit verified on `origin/main`, or
explicitly left open or quarantined with the reason, and `main` CI is green after the last merge;
* every review comment is answered and resolved;
* the campaign ledger summarizes the baseline clusters, root causes, fixes, before/after evidence,
capture provenance, product commits, merged PRs, and final matrix status;
* `self-heal` has applied durable campaign and harness lessons in the homes future runs will read —
and any newly earned capture-path mechanism (a way a capture produces a plausible wrong number or
a plausible absence) lands in `docs/PROFILING-CAMPAIGNS.md` specifically, the catalogue every
future capture session must read. The 2026-08 campaign routed such mechanisms there same-day for
ten days and then stopped in its final wave, losing two; the catalogue is a named self-heal
target, not an optional home.
**Every completion or merge-ready claim rests on its authoritative check, re-run after the last
change the claim covers — never on a proxy.** Name the check beside the claim: matrix state cites
the freshly regenerated outputs, a merge cites its commit on `origin/main`, "CI green" cites the run
for the exact head SHA, "rig left as found" cites the current holder, not the port. A claim whose
own verification failed — or ran before the last change — is withdrawn, not softened. The 2026-08
corpus has five completion claims resting on proxies (a stack membership asserted 23 s after its own
check failed; a nine-PR stack called merge-ready on per-PR CI alone), every one with the
authoritative check available and cheaper than the retraction.
Ask the user only for genuinely human-only device interaction, missing authorization, or a choice
that materially changes product behavior or campaign scope. Otherwise operate autonomously until the
completion gate or a control message is satisfied.