Skip to content
Back to skills

Run Rq

BSecurity

Use this skill to drive a research question (RQ) end-to-end: validate the RQ README, generate a fill batch-plan, start the Docker batch in the background, monitor progress, run aggregation, and propose findings updates. Trigger when the user says "RQ-N voranbringen", "run-rq", "fill RQ-N", "run RQ-N", "Forschungsfrage N starten", or names a specific RQ-N directory in research/.

  • 12 stars
  • 0 votes
  • 0 copies
  • 0 views
  • Added October 2, 2026
code-qualitypythongobashdockerrefactoring

Works with

  • cursor
  • terminal
  • cli

Security analysis

B85/100
  • highPerforms destructive filesystem operations

Pro shows the line behind each finding and how to fix it

Scanned October 2, 2026

npx -y skills add marcoemrich/agentic_coding_lab --skill run-rq --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Run Rq?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Run Rq
[![Security: B — Skills Directory](https://www.skillsdirectory.com/api/skills/marcoemrich-run-rq/badge)](https://www.skillsdirectory.com/skills/marcoemrich-run-rq)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: run-rq
description: |
  Use this skill to drive a research question (RQ) end-to-end: validate the
  RQ README, generate a fill batch-plan, start the Docker batch in the
  background, monitor progress, run aggregation, and propose findings updates.
  Trigger when the user says "RQ-N voranbringen", "run-rq", "fill RQ-N",
  "run RQ-N", "Forschungsfrage N starten", or names a specific RQ-N directory
  in research/.
---

# Skill: run-rq

End-to-end orchestration for advancing a single research question (RQ) in this lab repo. **Pure orchestration** — every operation calls existing repo scripts; no new Python or Bash code is written.

## Argument

- `RQ-N` (e.g. `RQ-model-quality`) or a direct path to an RQ dir.
- If not given: infer from the last user turn, otherwise ask back ("Which RQ? e.g. RQ-model-quality").

## Repo conventions (from the top-level `README.md` and memory)

- RQ dirs live in four subtrees: `research/questions-claude/<chapter>-*/` (Claude-Code RQs), `research/questions-opencode/<chapter>-*/` (OpenCode RQs), `research/questions-cross/<chapter>-*/` (cross-harness RQs), and `research/workflow-dev/<chapter>-*/` (workflow evolution). The `<chapter>` prefix (e.g. `2.6`) is an **ordering label, not an id** — the stable identity is the frontmatter `id:` (e.g. `RQ-lean`). Each RQ dir holds `README.md`, `findings.md`, `runs.csv`, `summary.md`.
- **Resolving an `RQ-<slug>` id to a path** (the dir name carries a chapter number, not the id): grep all subtrees for the frontmatter `id:`. Anchor with `^id:` and a trailing boundary so the whole slug must match exactly (no slug is a prefix of another, so an exact-line match is unambiguous):
  ```bash
  RQ_DIR=$(grep -rlE "^id:[[:space:]]*RQ-model-quality[[:space:]]*$" \
             research/questions-claude/*/README.md \
             research/questions-opencode/*/README.md \
             research/questions-cross/*/README.md \
             research/workflow-dev/*/README.md \
           2>/dev/null | head -1 | xargs -r dirname)
  ```
  On no match → ask the user. On multiple → take the first and inform the user. Pass `"$RQ_DIR"` to all scripts below (they accept any path and write outputs to the dir).
- Mandatory frontmatter fields: `id, question, factors, controls, outcomes, min_replicates`. There is no `status` field — whether the RQ needs work is read off the data (`experiments/rq-status.py`, rules in `experiments/rq_facts.py`). Only an RQ that ends without full data gets `closed: "<reason>"`, and only on the user's decision.
- Methodology constraint: `baseline-oneshot-*` / `baseline-iterative-*` only with `prompt: prose`; every other arm with all three styles. If `factors.workflow_x_prompt` exists, no additional `factors.workflow` / `controls.workflow` is allowed.
- Active katas: `claim-office`, `game-of-life`, `sphinx-score`, `game-of-life-cli`, `claim-office-lite`, `mars-rover`. `controls.kata_base` must be from this set. Each has the three prompt variants (`-prose`, `-user-story`, `-example-mapping`); all but `mars-rover` also have a `<basename>-verification/` suite, so `verification_pct` is available there.
  - The list is not a ranking, but the pool is lopsided in practice: `claim-office` and `game-of-life` carry the bulk of the runs, `sphinx-score` is the established small quality kata, and `mars-rover` is near-unused. Prefer a kata that already has runs in neighbouring RQs — a fill on a fresh kata has no reference cells to compare against.
  - **Check the actual kata dir before rejecting an RQ on this list.** The list is hand-maintained and has lagged behind the repo before (`sphinx-score` was in use in three RQs while still missing here). `ls experiments/katas/` is the authority; this line is a convenience copy.
- Model IDs are **lab-variant IDs** (`opus-4-7`, `opus-4-7-no-thinking`, `opus-4-6-portkey`, `opus-4-6-portkey-no-thinking`, `sonnet-4-6`, `sonnet-4-6-no-thinking`, `sonnet-4-6-portkey`, `sonnet-4-6-portkey-no-thinking`, `haiku-4-5`, `haiku-4-5-no-thinking`, `haiku-4-5-portkey`, `haiku-4-5-portkey-no-thinking`). The `-portkey` suffix marks models routed via the Portkey gateway.
- Aggregation is query-based: ALL runs in `experiments/runs/` matching the selector query count — regardless of which batch produced them.
- Batch plan is idempotent: counts existing matches and only fills missing replicates up to `min_replicates`.

## Phases

Run sequentially. On errors in any phase, **stop and ask the user**, do not skip ahead.

---

### Phase 1 — Validate

1. Resolve the RQ path into `$RQ_DIR` via the id-grep in "Repo conventions" above. On no match, ask the user; on multiple, take the first and inform the user.
2. Read `$RQ_DIR/README.md`.
3. Parse the frontmatter block (between the first two `---` lines). Check mandatory fields: `id, question, factors, controls, outcomes, min_replicates`. Missing fields → abort phase, inform user.
4. Check methodology constraints:
   - If `factors.workflow_x_prompt` is set: no additional `factors.workflow` and no `controls.workflow` may be set.
   - In every `workflow_x_prompt` entry: if `workflow ∈ {baseline-oneshot-v1-cc, baseline-iterative-v1-cc}`, then `prompt == prose` is required.
   - `controls.kata_base` is one of the active katas listed under "Repo conventions" above. Verify against `ls experiments/katas/` rather than against the list alone — the list is a convenience copy and has lagged the repo before.
   - Model values (in `controls.model` and/or `factors.model`) must appear in the lab-variant table.
5. Read `findings.md` (needed in phase 6 as the existing baseline).
6. Target computation: from `factors` × `controls` derive the cell count (every factor multiplies; paired factors like `workflow_x_prompt` count as a single factor with `len(pairing)` values). Target runs = cells × `min_replicates`. Report this number to the user.
7. **Portkey routing detection**: scan model values (in `controls.model` plus every `factors.model` entry) for a `-portkey` suffix. If any match:
   - Check `~/.claude.portkey/` directory exists. If missing: STOP, instruct the user to follow `experiments/docker/claude-config-portkey.README.md`. Do **not** start the batch.
   - Note: `batch.sh` auto-detects Portkey models from the plan JSON and sets `CLAUDE_CONFIG_DIR` automatically. No manual env-var needed in phase 3.

Output to user (compact):
```
RQ-N validated: <id>, <#cells> cells × min_replicates=<n> = <target> target runs.
cells at min_replicates: <full>/<declared> · findings: <n> · runs newer than findings: <n>
[Portkey routing required — using ~/.claude.portkey/ profile]   ← only if portkey_required
```

---

### Phase 2 — Plan

1. Run dry-run:
   ```
   experiments/batch-plan-from-rq.py "$RQ_DIR" --dry-run
   ```
2. Output contains the count of missing runs. Inform user:
   ```
   Cells: X, missing runs: Y → would write experiments/batch-plans/rq-{n}-fill.json
   ```
3. If `Y == 0`: no new runs needed — jump straight to phase 5.
4. Otherwise: get user confirmation ("Write the plan with Y runs now?"). Only proceed after explicit "yes".
5. Write the plan without `--dry-run`:
   ```
   experiments/batch-plan-from-rq.py "$RQ_DIR"
   ```
6. Briefly inspect the generated plan (Read on `experiments/batch-plans/rq-{n}-fill.json`) and summarize the cell distribution (`{kata, workflow, model}` frequencies) to the user.

---

### Phase 3 — Run

1. **Pre-check**:
   - `docker ps --filter name=docker-batch-run --format '{{.Names}}'` — if a container is running: STOP and ask the user whether the existing batch should finish first.
   - If `experiments/docker/batch.log` from an earlier run exists and is >0 bytes: ask the user whether to back it up as `experiments/docker/batch.<plan>.log` (`mv`, no deletion).
2. **User confirmation** before start: "Start batch `rq-{n}-fill` with Y runs in the background? Expected wallclock ≈ Y × 6 min ≈ Z min." (6 min/run per memory, smart-subset experience.) For Portkey RQs (flag from phase 1.7), append: "via Portkey gateway (auto-detected)".
3. Start with `--detach` (backgrounds the batch and disowns the process so the terminal can be closed safely):
   ```bash
   cd experiments/docker && ./batch.sh rq-{n}-fill --detach
   ```
   Portkey routing is **auto-detected** by `batch.sh`: it scans the plan JSON for `-portkey` model names and sets `CLAUDE_CONFIG_DIR=~/.claude.portkey` automatically. No manual env-var override needed. Sharding is on by default (5 shards, round-robin); override with `--shards N` when a plan is small or should run in one container:
   ```bash
   cd experiments/docker && ./batch.sh rq-{n}-fill --shards 10 --detach   # large fill
   cd experiments/docker && ./batch.sh rq-{n}-fill --shards 1 --detach    # single container
   ```
4. After a few seconds, determine the container name via `docker ps --filter name=docker-batch-run --format '{{.Names}}'` and report it to the user.

---

### Phase 4 — Monitor

1. Polling loop with `experiments/docker/watch-batch.sh rq-{n}-fill`:
   - First 5 min: snapshot every 60 s.
   - After that: every 5 min.
   - On every poll, report the counter `[N/total]` and container status.
2. Termination conditions:
   - `Container STOPPED` AND counter = `[total/total]` → success, continue to phase 5.
   - `Container STOPPED` AND counter < total → resume (phase 4b).
   - User signals abort → `docker stop <container>` → resume (phase 4b).
3. **Important**: Take memory patterns seriously:
   - `\b429\b` with `claude_exit != 0` = real rate limit; backup-warning with `backup.<ms>.json` (which can incidentally contain `429`) is to be IGNORED.
   - "Claude configuration file not found at: …" at container start is harmless.
   - Do not panic-retry when `watch-batch.sh` is once slow to respond.
4. **Mid-execution cleanup after stop**: the youngest run dir in `experiments/runs/` that lacks `analysis-report.md` OR `transcript.jsonl` was interrupted mid-execution. Ask the user: "Delete interrupted run dir `<run-dir>` (no analysis-report.md)?" Only delete with `rm -rf` after explicit "yes".

5. **Reading partial results while the batch runs** — optional, but if you look at
   metrics mid-batch, these three rules are binding. All three failed in one session
   (2026-08-10) and produced three wrong statements to the user in a row.

   **a) Never quote a quality metric without `verification_pct` in the same query.**
   A run that failed external verification produces *excellent-looking* quality
   numbers — few functions, low complexity, no smells — because it barely implemented
   anything. The correctness-gating rule of phase 6 exists for exactly this, but it
   only helps if the correctness column is in front of you. Query it together, always:
   ```bash
   jq -r 'select(.run_status.exit_reason=="ok") |
     "vpct=\(.final_metrics.verification_pct) cog=\(.final_metrics.cognitive_max) …"' */metrics.json
   ```
   Real case: `cognitive_max` 1 and `cc_avg_loc_per_function` 4.67 were reported as
   "beats the best cell in the field" — the run had `verification_pct = 0`.

   **b) A run is only readable once it is finished.** `metrics.json` exists from the
   start and is filled at the end; a running run returns `null` for every metric and
   an absent `exit_reason`. Filter on `select(.run_status.exit_reason=="ok")` rather
   than on file existence, and never mix finished and running runs in one listing —
   `ls -dt | head -1` regularly hits a running one. (Same trap as the cursor-harness
   note in memory: judge runs only after `experiment-done.txt`.)

   **c) Do not interpret a cell at n=1.** Report partial values as an observation
   ("first run of cell X shows …"), never as a rank statement, a hypothesis
   confirmation, or a comparison against a reference cell that has full n. The
   variance across replicates in this lab routinely exceeds the between-cell
   differences being measured — `refactorings_applied` σ 17.4 at n=5 in
   RQ-architecture-axis-sol-pi F-1.4 is a documented example.

   Also watch the glob when spot-checking: `*hybrid-v6*opus-5*` matches pi and cursor runs
   from other RQs. Anchor workflow and model exactly — `*_exact-hybrid-v6-lab-split-cc_opus-5-no-thinking*`.

#### Phase 4b — Resume

1. Generate a resume plan:
   ```
   experiments/docker/resume-plan.sh rq-{n}-fill
   ```
   → writes `/tmp/rq-{n}-fill-resume.json`.
2. Show the size to the user (`jq '.runs | length' /tmp/rq-{n}-fill-resume.json`).
3. Get user confirmation: "Restart with `<m>` remaining runs?"
4. After "yes" — resume with `--detach` (Portkey auto-detected from plan content):
   ```bash
   cd experiments/docker && ./batch.sh /tmp/rq-{n}-fill-resume.json --detach
   ```
   Back to phase 4.

---

### Phase 5 — Aggregate

1. **Pipeline sanity check before aggregating** — pipeline bugs masquerade as research findings. Spot-check 2–3 of the matched runs (covering the workflow × model cells most central to the RQ) for divergence between `run.log` and `metrics.json`:
   - In `run.log` the agent typically reports a final test status (e.g. "All N tests pass", "experiment-done.txt written"). Compare this against `final_metrics.tests_passing` in `metrics.json`. If the agent reports green but `tests_passing` is `false`, that's a pipeline issue, not a workflow effect.
   - Grep `analysis-report.md` for known infrastructure failure patterns: `IGNORED_BUILDS`, `approve-builds`, `corepack`, `ENOENT`, `Cannot find module`, `tsc.*error`. These typically come from container/tooling drift (e.g. pnpm version bumps, missing deps), not from agent code.
   - For CLI-katas (`<basename>-verification/` exists): check `cli_built` in `metrics.json`. If `cli_built: false` for a run whose `src/cli.ts` exists and runs manually (`pnpm exec tsx src/cli.ts < scenario.input.json`), the verification stage misfired — re-run `analyze-run.sh <run_dir>` on the host (with absolute path).
   - If any of these checks fail: stop, report the suspected pipeline bug to the user, do NOT aggregate. Fixing buggy data after a finding has been drawn from it is much more expensive than spending two minutes on the spot-check.
2. **Post-processing before aggregation** — some metrics are not filled by `analyze-run.sh` and must be computed separately. Skipping a step here does not error; it silently produces a zero or a null that reads like a research result.
   - **`cost_usd` — always, for any RQ whose `outcomes` contain it.** pi/Requesty runs land with `cost_usd = 0` because Requesty reports no inline cost (`usage: null`); the value comes from token × list price. Run it before aggregating:
     ```
     experiments/compute-cost.py "$RQ_DIR"
     ```
     Idempotent — existing values are recomputed, so it is safe to run on an RQ whose older runs already carry costs.

     **Ordering caveat — this step reads `runs.csv`, which phase 5.3 writes.** On an RQ that was just filled (or that you reached with 0 missing runs), `runs.csv` still describes the *previous* run set, so the script costs a subset and reports success for it. The script now refuses a stale csv and tells you to aggregate first; if you hit that, run `aggregate-by-query.py` once, then this script, then `aggregate-by-query.py` again so the fresh costs reach `summary.md`. Two verifications afterwards, not one: no cell may show $0.00 at non-zero `total_tokens` (impossible combination), **and** every cell present in the coverage table must also appear in the `cost_usd` pivot. A cell missing from that pivot entirely is the more common failure and reads like "metric not applicable" rather than like an error.
   - **`mutation_score` — only when it appears in `outcomes:`.** Expensive (minutes per run), opt-in per RQ, and only computed for `tests_passing = true`:
     ```
     experiments/compute-mutation-score.py "$RQ_DIR"
     ```
3. Invoke:
   ```
   experiments/aggregate-by-query.py "$RQ_DIR"
   ```
4. Expected outputs:
   - `$RQ_DIR/runs.csv` (one line per matched run)
   - `$RQ_DIR/summary.md` (per-cell pivots for each `outcome`)
5. Read `summary.md` in full and summarize to the user — show the per-cell pivot tables individually.
6. Sanity check: does every cell have ≥ `min_replicates`? If not: warn and offer to jump back to phase 2 (additional runs).
7. **Plausibility cross-check before phase 6** — if any cell value contradicts a previously-stable finding by a large margin (e.g. a workflow that was 100% green is suddenly 0%), do NOT treat that as a new finding without first running the spot-check from step 1 against that exact cell. A "the subagents arm is suddenly broken on game-of-life" type observation is more often a pipeline regression than a real shift.

---

### Phase 6 — Findings (write-first)

**Write directly to `findings.md`, then notify the user to review.** Markdown tables and trophy assignments are much easier to evaluate as rendered output than as a chat proposal; reverting is cheap (it's only markdown). After writing, send one line (in the user's language), e.g. "written — please review",.

Exception: **deletions** of existing findings still require explicit user confirmation before the `Edit` — losing a documented finding is more expensive than re-reading a fresh write.

`findings.md` shows **only the current state**. No status tags (`✅ stabil` / `⚠️ bedingt` / `🚫 offen` / `❌ widerlegt`), no comparisons with archive snapshots or older studies, no "previously X, corrected" hints in the prose. Header form: `## F-x.y — title` (no `· …` suffix).

**Overview table**: `findings.md` starts with a `## Overview` section containing a pivot table of the primary outcome across all factor levels (all models, all prompt styles, etc.) — before the individual `F-x.y` blocks. This table gives readers the full picture at a glance; individual findings then zoom in on specific effects. Update this table whenever findings are added or updated.

**Trophy convention (🏆) in overview tables**: append 🏆 to the best value per outcome row. Conventions:

- **Comparability first — a 🏆 needs an actual contest.** Before awarding anything, ask what varies across the columns of that row. Trophies are only meaningful when the cells differ in the factor under study and are otherwise alike. If a row spans cells that differ in *task size* rather than in the factor, drop the trophies for that table entirely and say so in the caveat block. The clearest case is an RQ with `kata_base` as a factor: `sphinx-score` (~183 Code Mass, ~12 cycles) against `claim-office` (~997, ~46) — the small kata always "wins" on cost, complexity and Code Mass (APP), which measures the kata, not the work. Same for any cross-kata cost or shape row. Options in order of preference: split the table so each row compares within one kata, restrict trophies to the rows that carry a real contrast, or omit them and label the table as context. This test comes **before** correctness-gating: gating narrows an existing field of competitors, it never creates one.
- **The test-list boundary is a comparability failure, not a close call.** Before
  awarding a trophy on `tdd_discipline`, `tdd_discipline_test_first`,
  `tdd_discipline_closure`, `test_first_rate`, `skip_events`, `cycles_total` or
  `chain_deviations`, check whether the row's cells differ in whether a test list
  is written up front. A `Skip` is compliance in a test-list workflow and a
  start-over condition in a strict-red one, so those columns then measure the
  architecture. Split the table by group and award within each, as
  RQ-tdd-workflow-comparison-opus55 does — never one trophy across the boundary.
  `aggregate-by-query.py` prints a warning naming both groups when an RQ is in
  this situation; treat it as binding. `red_batch_max` is the column that does
  compare everywhere, and on the measured field it orders the cells the opposite
  way from the score.
- The direction is metric-dependent — note it in the column header or row label (`smell_total` etc. → "lower = better"; `refactorings_applied`, `predictions_correct_rate` → "higher = better"). Don't assume.
- **Check the trophy against its own row after writing.** The winner must be the best value *in that row* under the stated direction. Gating rules constrain which cells are eligible, but they never move the trophy onto a worse value: if the only eligible cell is not the row's best, that row gets no trophy. Two failure modes to look for — a 🏆 on a higher number in a "lower = better" row, and a 🏆 in a row where every cell is identical (no contest, so no winner).
- Use 🏆 only where there is a meaningful winner. If the spread is below 1 σ and the framing is "no effect", award 🏆 to all near-tied values (or to none if the table message is "indistinguishable") — don't fabricate a winner from rounding noise.
- Multiple 🏆 are fine for ties. Three 🏆 across a row signal "no effect", which is itself a useful reading aid.
- Always bold the winner value too — 🏆 is in addition to, not instead of, the bold.
- Trophies belong only in human-facing research documents (`findings.md`, archive snapshots). Workflow files (`experiments/workflows/**/*.md`) stay emoji-free per the RQ-emoji / CLAUDE.md convention.
- **Correctness-gating** for code-quality and efficiency metrics: when the RQ has a correctness outcome (typically `verification_pct`), trophies for quality/efficiency metrics (`smell_*`, `cognitive_*`, `mccabe_*`, `cc_*`, `duration_seconds`, `total_tokens`, `cost_usd`, also derived $/perfect-result ratios) go **only** to cells with mean `verification_pct ≥ 0.90` — cells, not just models: the same applies when the factor is a prompt style or a kata. A cell that scores low on complexity / cost / duration but failed the verification is showing a stub, an abort artifact, or an implementation of a smaller wrong rule set — not parsimony or speed. Awarding a 🏆 there is misleading.

  The threshold is 0.90 rather than a perfect 1.0 because several katas have a case that no model solves — on claim-office `14-family-steinheim` fails deterministically for every cell (RQ-1.19 F-1.19.4), which caps the achievable mean at 14/15 = 0.93. A literal 1.0 gate would disqualify every cell in such an RQ and leave the table without a single winner, which says nothing about the models and defeats the point of the overview. 0.90 admits cells that solve essentially the whole spec while still excluding stubs and aborts.

  **A cell above the threshold can still be disqualified for a specific metric** when its leading value comes from work it did not do. Judge the failure mode, not only the score: a cell that drops an entire class of scenarios has a lower Code Mass (APP) because of missing coverage, not parsimony. Mark such a leading value in parentheses and award the trophy to the best cell that did the full work — e.g. `(659.0)` with no 🏆, the trophy going to the next-lowest eligible cell. State the reason in the caveat block. State the rule explicitly in the table's caveat block when it applies. Pure correctness metrics (`verification_pct` mean / std) are not gated. RQs without a correctness outcome (pure code-quality studies on game-of-life) are also not gated. Gating is a filter on an existing contest, not a licence to award: if it leaves a single eligible cell in a row whose columns were never comparable, the answer is no trophy, not a trophy for the survivor.

1. Diff sources:
   - **Existing**: `findings.md` from phase 1.
   - **New**: `summary.md` from phase 5.
2. Three possible actions per effect:
   - **New finding**: cell/factor group with Δ ≥ 1σ over the other groups AND the effect is not yet covered in `findings.md` → new `F-{N}.{M+1}` block (M = highest existing finding number). Write directly.
   - **Update**: an existing finding covers the same effect, but cell values or interpretation have shifted → `Edit` the existing block directly. Rewrite table and rationale, **without** old/new diff, **without** "previously X", **without** reference to archive snapshots.
   - **Deletion**: data contradicts the finding → ask the user first ("Finding F-x.y is contradicted by new data — remove it?"), then remove the block including its `---` separator on confirmation. Do not mark it as "refuted".
3. Data gap: if an effect is suspected but coverage is too small for `n ≥ min_replicates` → note in `todos_and_ideas/1-IN.md` as a bullet with a concrete re-check target. **Do not** create as a finding in `findings.md`.
4. Format per block: statement / data-base table / rationale. Header `## F-x.y — title` with no suffix. The namespace before the last dot may carry dots itself (`F-4.4.1`, `F-1.12.5` — chapter-numbered ids are in use and valid). The em-dash `—` is load-bearing: `generate-snapshot-skeleton.py` parses headers on it and silently skips any it cannot match, which then reads as "no findings documented" in the snapshot for an RQ that has a full findings.md.
   **Glossary discipline**: terms like `code_mass`, `cc_loc`, `cc_longest_function`, `smell_total`, `verification_pct` are to be used only in the form from the glossary in the top-level `README.md` ("Code Mass (APP)", "Production LoC", "Smell Total", "Correctness (external)") or directly via the metric ID in backticks. Synonyms like "Code-Volumen", "Code-Gesamtvolumen", "LoC-Größe" are forbidden — they are ambiguous or collide with established definitions (APP). The three complexity metrics carry **no** prose name at all: write `cc_longest_function`, `cognitive_max`, `cognitive_avg` or `mccabe_max` in backticks. "Complexity Peak", "Spitzen-Komplexität", "Cognitive Complexity peak" and "McCabe peak" are forbidden; the first had been used for three different metrics at once. Before writing, read the glossary once and check every term used in the block against the table.
5. **Number formatting in `findings.md`:** large counts (typically `total_tokens`, `subagent_token_total`, cache stats) get the `M`-suffix (millions) or `k`-suffix (thousands), not scientific notation. `summary.md` keeps the pandas-default `e+07` form — only `findings.md` gets reformatted. Examples: `44.4 M` (good), `4.44e+07` (bad in findings, fine in summary), `1023 ms` or `1.0 s` (good for duration). σ-Werte werden im selben Format dargestellt wie der Mean (`σ ≈ 5 M`, nicht `σ ≈ 5e+06`). Rationale: M/k ist auf einen Blick lesbar, e+07 zwingt zum Kopfrechnen.
6. **Verify every number you just wrote** before reporting. Do not skip this — a wrong
   number in a findings table looks authoritative and gets quoted onward. Pull the written
   rows back out and diff them against `summary.md`:
   ```bash
   grep -n "<cell-name>" "$RQ_DIR/findings.md" | grep "|"
   ```
   Check each value against the same metric in `summary.md` for **this** RQ.

   The failure mode this catches: when several RQs are edited in sequence, values from a
   different kata sit in context and get written from memory instead of looked up. Cell
   names and metric names are identical across RQs — only the ranges differ, and a foreign
   value is often plausible enough to survive proofreading. (Real case, 2026-08-04:
   `cc_longest_function = 15.0` from the game-of-life RQ landed in the claim-office RQ,
   where 21.4 was correct — wrong trophy, plus an interpretation sentence built on the
   wrong figure.)

   Two habits prevent it upstream: query metrics **completely** per RQ rather than
   selectively (the gap appears when a table column is missing from the fresh extract),
   and treat magnitude as a sanity anchor — claim-office is the large kata
   (`code_mass` ~666, `cycle_count` ~47), game-of-life the small one (~144, ~15).

7. After verifying, send one short line to the user (in their language), e.g. "written — please review". The user reviews the rendered markdown directly.

---

## Out of scope (deliberately NOT in the skill)

- The actual batch execution inside the container (`run-batch.sh` runs in the container).
- ESLint / smell detection (runs per run inside `analyze-run.sh`).
- Cross-RQ aggregation or creation of new RQs.
- Auto-commit/push of `runs.csv` / `summary.md` / `findings.md` — stays a user decision.
- Merging into `main`.

## Behavior on errors

- **Phase 1 fails**: do not continue; give a clear constraint hint (e.g. "baseline-oneshot-v1-cc with prompt=example-mapping violates the methodology constraint in the top-level README, section 'Methodology constraints'").
- **Phase 2/3 scripts with non-zero exit**: show output to the user; do NOT blindly retry.
- **Phase 4 loses the container**: show `docker ps -a`, then offer resume.
- **Phase 5 produces an empty `runs.csv`**: check whether `experiments/runs/` actually contains matching runs (selector too narrow?). Inform the user.

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…