Clean-room test of the Spur server install by planting two coding agents — cursor-agent and Claude Code — on a persistent cloud VM that gets reset to a pre-install state each run; the harness never creates or deletes a VM. Cron-safe — runs end to end from the skill name alone, opens a PR for any doc fix, reports in-session. Reset the box over plain SSH, plant both agents by copying the operator's credentials (itest-only shortcut), run each on the README Install block verbatim, capture the tra...
Pro shows the line behind each finding and how to fix it
Scanned 10/2/2026
npx -y skills add ashugaev/spur --skill clean-install-test --agent claude-codeInstalls into .claude/skills of the current project.
Are you the author of Clean Install Test?
Add the live security badge to your README — it updates automatically with every re-scan.
[](https://www.skillsdirectory.com/skills/ashugaev-clean-install-test)More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.
---
name: clean-install-test
description: Clean-room test of the Spur server install by planting two coding agents — cursor-agent and Claude Code — on a persistent cloud VM that gets reset to a pre-install state each run; the harness never creates or deletes a VM. Cron-safe — runs end to end from the skill name alone, opens a PR for any doc fix, reports in-session. Reset the box over plain SSH, plant both agents by copying the operator's credentials (itest-only shortcut), run each on the README Install block verbatim, capture the transcripts, turn friction into doc fixes, verify services, compare the two agents. Same box owns the source-install deploy path — use on any change to `scripts/main-deploy.sh`, `deploy/*.service`, or that deploy path, and to verify a deploy end to end on the itest VM.
---
CLEAN INSTALL TEST
A role-play: reset a persistent VM to a pre-install state, plant two coding agents on it, act as a non-coding user who hands each agent only the Spur docs, zero hints. Whatever a fresh agent can't do from the docs alone is doc friction to fix — the agent does the install, you observe, analyze, evolve the docs. Goal: a fresh agent, given only the docs, installs Spur single-shot up to the identity gates (agent login, Tailscale login), collected into a final user TODO, not hacked around. Iterate the docs until that holds for both agents.
Runs with no prompt beyond the skill name and no human present — see AUTONOMOUS MODE.
GROUND RULES
- Never create a VM. Never delete a VM. One box stays up permanently; reset it in place before each run. Lifecycle: read the recipe -> reset the box -> plant both agents -> run the test on each -> analyze -> verify services -> evolve docs -> report. Provisioning a replacement box is a one-time exception on the user's explicit instruction only, never part of this cycle — see APPENDIX at the end.
- Ubuntu 24.04 LTS. e2-small is the floor machine, not e2-micro: the planted agent runs ON the box and competes with the install for CPU. e2-micro wedged mid-run on a repeat test — CPU credit exhaustion, not RAM (peaked 686 MB of 955 MB on the pass run) — sshd stopped answering, serial console showed `systemd-networkd: Could not set DHCPv4 address: Connection timed out`, a stop/start did not recover it within 5 min. Resized to e2-small 2026-08-23: 1961 MB RAM, 2 vCPU. npm bundle ships the web UI prebuilt, no on-box build. The source-install mode below runs a full `pnpm install` plus a Next build on the box — that path sets the floor. Never open app ports to the internet — default firewall leaves only SSH reachable, leave it.
- The external IP is a reserved static address, reserved 2026-08-23. Stop/start keeps it, so a resize or a recovery stop/start is allowed. Never release the address reservation.
- No hard hacking: a planted agent never bypasses an identity/auth step; lacking the user's own account it records a TODO and moves on — never hint it past friction, fix the doc instead. Copying credentials onto the VM (step 3) is an ITEST-ONLY harness shortcut for an unattended run: never outside itest, never between real hosts. Secrets and credentials go to the VM only, piped over SSH stdin, never echoed to logs or chat.
- A run that finds zero friction is a pass. Report it and stop — never invent a doc edit to justify the cycle.
0 PREREQS (LOCAL BOX)
SSH needs no cloud CLI once the key is baked in — a permanent SSH keypair, generate once if missing: `ssh-keygen -t ed25519 -f <keypath> -N ''`. A host-local recipe file (e.g. `~/.spur/itest-conn.md`) holding the box's project/zone/key paths and its current name/IP — read it first, the persistent box already exists. `~/.config/cursor/auth.json` and `~/.claude/.credentials.json` on the operator's box — itest-only, both planted agents run on them.
1 CONNECT (CLOUD-CLI-INDEPENDENT)
ssh -i <keypath> -o StrictHostKeyChecking=no -o UserKnownHostsFile=/dev/null <user>@<IP> '<cmd>'
The box's name and IP live only in the host-local recipe. That entry is a cache; the cloud is truth. SSH does not answer: the box is not wedged and a replacement is not the fix. Re-verify against the cloud first, then reconcile the recipe:
gcloud compute instances list --filter='name~itest' --format='value(name,status,EXTERNAL_IP)'
Measured 2026-08-23: the recipe named an instance absent from the project, and filed the live box — 6 days uptime, different IP — under a "dead box, awaiting deletion" heading. Three SSH attempts at a 30s connect timeout all timed out. Reading the recipe alone points at bootstrapping a second VM for nothing.
2 RESET THE BOX (EVERY RUN)
Upload the committed reset script over stdin so the box always runs the version in this repo, then run it and read its printed state table before trusting the box is clean:
cat tests/itest/reset-vm.sh | ssh ... 'cat > ~/itest-reset.sh && chmod +x ~/itest-reset.sh'
ssh ... 'bash ~/itest-reset.sh'
`tests/itest/reset-vm.sh` never removes `~/.itest-harness`, `~/.claude/.credentials.json`, or `~/.config/cursor/auth.json` — the Claude harness node and both agents' credentials survive a reset by design. It does clear system-scope `spur-*.service` units and `/etc/spur`, which the source-install deploy mode leaves behind: a stale system daemon holds 4310 across a reset and the next npm run reads it as its own. Same source: `~/spur` and `~/spur-mirror`, plus `~/projects` from a previous run's smoke project — a tested agent that finds that checkout's maintainer `spur.yaml` spends turns deciding whether to connect live GitHub and Telegram sources, friction the harness created. It also wipes `~/.claude/skills` and `~/.codex` wholesale (never `~/.claude` itself) so a persistent box re-exercises the fresh-host branch of host-skills install every run. Every agent CLI the doc installs conditionally (`claude`, `codex`, `opencode`) is removed too: a leftover copy silences its `command -v ... ||` line, and the run reads as a pass while never exercising it — measured 2026-09-07, an `opencode` from 2026-08-25 survived every reset. Its state table also reports `claude`, `codex`, and `opencode` — `present` on any doc-installed agent means stop and fix the box, not run the test. It ends with `port-4310`, `harness`, `harness-creds`, `agent-skills`, and `source-clone` — `BUSY`, `MISSING`, or `leftover` there means the same.
3 PLANT THE AGENTS (ITEST-ONLY HARNESS) — BOTH REQUIRED, OFF THE TESTED PATH
Both cursor-agent and Claude Code get planted on the box; both runs are required, not primary/fallback — see 4 RUN THE TEST.
cursor-agent is self-contained, needs no node, and the box stays node-free until the tested agent installs node itself — that install is part of what the test measures.
ssh ... 'curl https://cursor.com/install -fsS | bash' # lands ~/.local/bin/cursor-agent
cat ~/.config/cursor/auth.json | ssh ... 'umask 077; mkdir -p ~/.config/cursor; cat > ~/.config/cursor/auth.json'
ssh ... '~/.local/bin/cursor-agent --list-models' # pick a weak model; this run used gemini-3.7-flash-high
ssh ... '~/.local/bin/cursor-agent -p "Reply with exactly one word: authok" --model gemini-3.7-flash-high --force --output-format text'
`--force` is required on every `-p` run — without it the run dies on a directory-trust prompt ("Pass --trust, --yolo, or -f"), not an auth failure. Caveat: cursor pulls the operator's account-level user rules into the planted agent's context — the run is not a pristine blank agent.
Claude Code ships a native binary: `~/.itest-harness/bin/claude` resolves to `lib/node_modules/@anthropic-ai/claude-code/bin/claude.exe`, needs no node at run time. A `node .../cli.js` invocation fails MODULE_NOT_FOUND — never use it. Installing it needs node once — install a private harness node off the tested PATH and use its npm:
mkdir -p ~/.itest-harness
curl -fsSL https://nodejs.org/dist/v22.14.0/node-v22.14.0-linux-x64.tar.xz | tar xJ -C ~/.itest-harness --strip-components=1
PATH=$HOME/.itest-harness/bin:$PATH npm install -g --prefix $HOME/.itest-harness @anthropic-ai/claude-code
cat ~/.claude/.credentials.json | ssh ... 'umask 077; mkdir -p ~/.claude; cat > ~/.claude/.credentials.json'
ssh ... 'PATH=$HOME/.itest-harness/bin:$PATH node -e "const f=require(\"os\").homedir()+\"/.claude/settings.json\";const fs=require(\"fs\");const d=fs.existsSync(f)?JSON.parse(fs.readFileSync(f)):{};d.skipDangerousModePermissionPrompt=true;d.skipAutoPermissionPrompt=true;fs.writeFileSync(f,JSON.stringify(d,null,2))"'
Launch with a sanitized environment so its child shells see the node-free box under test, never the harness node. Read the prompt from a file on the box (write it first — see 4 RUN THE TEST) instead of an ssh-local `"$PROMPT"`: the whole remote command sits inside one pair of local single quotes, so a local shell variable never expands there, it ships as its own three literal characters and the agent launches with an empty prompt.
ssh ... 'env -i HOME=$HOME USER=$USER TERM=dumb PATH=/usr/local/sbin:/usr/local/bin:/usr/sbin:/usr/bin:/sbin:/bin:$HOME/.local/bin ~/.itest-harness/bin/claude -p "$(cat /tmp/prompt.txt)" --dangerously-skip-permissions --output-format stream-json --verbose'
Verified: a planted Claude launched this way runs `command -v node` and gets NO-NODE. Same itest-only credential caveat as cursor-agent applies.
`env -i` also drops `XDG_RUNTIME_DIR`, so the agent's first `spur init` exits 1 with `npm-init: user systemd is unavailable`. Harness artifact, not doc friction — a plain `ssh host '<cmd>'` shell gets that variable from pam_systemd, so a real operator never sees it. An agent that exports `XDG_RUNTIME_DIR=/run/user/$(id -u)` and re-runs still counts as single-shot.
4 RUN THE TEST — README INSTALL BLOCK, SINGLE-SHOT, NO HINTS
Run twice per cycle, once per agent, with a full reset (step 2) between the two runs — a box dirtied by one agent's run is not a clean box for the other's.
The prompt is the Install block of the repo README verbatim — that block is the artifact under test, not a prompt the runner writes. It tells the agent to fetch `raw.githubusercontent.com/<owner>/spur/<ref>/docs/install-from-npm.md`, never a `github.com/.../blob/...` URL — that form returns an empty document under cursor's webFetch. `<ref>` is `main` by default; point it at the PR branch to test an unmerged doc fix. No docs are staged on the VM — the agent fetches this one file over HTTPS; do not tar the repo docs onto the box, the prompt never reads them.
Write the prompt to a file on the box AFTER each reset, one ssh call, no nested quoting to get wrong — the same file backs both agents' launches. `reset-vm.sh` deletes `/tmp/prompt.txt` as a run artifact, and `$(cat)` on the missing file launches the agent with an empty prompt:
ssh ... "cat > /tmp/prompt.txt" <<'EOP'
<the README Install block, verbatim>
EOP
Launch detached from `~` (no CLAUDE.md there), stream-json, poll a done-file, prompt read from the file with `$(cat ...)` — never a local `"$PROMPT"` inside the single-quoted remote command, that never expands and ships empty. The launch ssh call can hang even after the remote process has detached — never wait on it, verify with a separate ssh call instead:
ssh ... 'nohup bash -c "timeout 1800 ~/.local/bin/cursor-agent -p \"$(cat /tmp/prompt.txt)\" --model <m> --force --output-format stream-json </dev/null > /tmp/agent-run.jsonl 2>&1; echo \$? > /tmp/agent-run.done" >/dev/null 2>&1 &'
ssh ... 'pgrep -af cursor-agent; ls -l /tmp/agent-run.jsonl'
Long shell commands the planted cursor-agent runs go background inside cursor-agent itself: it gets `awaitToolCall` polls carrying a taskId, and the command's own output lands in `~/.cursor/projects/<slug>/terminals/<taskId>.txt` — check there when the jsonl shows a pending tool_call and nothing else.
5 ANALYZE THE TRANSCRIPT — FRICTION IS THE OUTPUT
cursor-agent stream-json event shapes, needed to parse a transcript at all:
{"type":"tool_call","subtype":"started"|"completed","tool_call":{"<name>ToolCall":{"args":{...},"result":{"success"|"failure":{...}}}}}
{"type":"assistant","message":{"content":[{"type":"text","text":...}]}}
{"type":"result","subtype":"success","duration_ms":...,"result":"<final message>"}
Smoke-test each agent's credentials before the RESET, not after — a reset costs minutes and proves nothing if an agent can't answer (`-p "Reply with exactly one word: authok"`). Both agents run on the operator's own accounts and share the operator's limits. cursor's planted `auth.json` goes stale on its own clock: `Authentication required. Please run 'agent login' first` from the PLANTED agent while the LOCAL one answers `authok` means the copy is stale — re-plant it (step 3) and smoke again. Same error from the local one means the operator must re-login, no VM-side fix exists. Claude's planted `.credentials.json` goes stale the same way, seen 2 runs in a row in 2026-09: `OAuth session expired and could not be refreshed` from the PLANTED claude while the LOCAL one answers `authok` — re-plant it (step 3) and smoke again. Local fails too: operator re-login. `Your team has reached its usage limit` names a date and no fix at all. Claude's planted run can die mid-install with rc=1 and `"api_error_status":429` / `You've hit your session limit` — the harness session burned the same account's window; reset and re-run once it resets, a full run is only ~350 s / 23-28 turns. A transcript under 1 KB with no tool_call event is a harness failure, not a doc failure — read the last line for a plain-text agent-side error (account quota, expired credentials, unavailable model), report that agent's run as blocked, never add it to the friction list. The other agent's run still stands on its own. One agent blocked: report the cycle as partial and name which agent produced evidence — never report a pass or invent friction to fill the gap. Same check before trusting any run: confirm the transcript's first user event carries the real prompt text, not an empty string — an empty prompt is a harness failure too.
Walk each tool_call with its result, each error, and the final result message. Look for steps the agent got wrong, retried, or did not infer from the docs (doc gap); anything it hard-blocked on versus correctly deferring to the user TODO; whether it chose the safe path (private/Tailscale, never public expose). Identity gates — agent login and `sudo tailscale up` — land in the final TODO as a pass, not friction, when reached cleanly with the user's action stated; real friction is anything the agent should have handled from the docs but didn't.
6 VERIFY SERVICES (INFRA LEVEL, NOT UI)
Confirm the agent's install works. Expected topology after `spur init`: two user units only.
systemctl --user is-active spur-daemon.service spur-web.service # both active
curl -sf -o /dev/null -w 'daemon %{http_code}\n' http://127.0.0.1:4310/sessions # 200
curl -sf -o /dev/null -w 'web %{http_code}\n' http://127.0.0.1:5555/ # 200
curl -s -o /dev/null -w 'ws %{http_code}\n' --max-time 5 \
-H 'Connection: Upgrade' -H 'Upgrade: websocket' \
-H "Sec-WebSocket-Key: $(head -c16 /dev/urandom | base64)" -H 'Sec-WebSocket-Version: 13' \
'http://127.0.0.1:5555/ws?session=none' # 101
ss -ltn
Pass = both units active, daemon 200, web 200, `/ws` upgrade 101, `ss -ltn` shows only 22, 127.0.0.1:4310, 127.0.0.1:5555, plus systemd-resolved — nothing else, and nothing answers on the app ports from the VM's public IP.
7 EVOLVE THE DOCS
For each real friction, edit the install doc minimally, then reset (step 2) and re-run both agents until each completes single-shot to the identity gates with a clean TODO. Fix the doc, re-test, repeat — never fix by hinting the agent.
8 REPORT
Per agent: single-shot or not, each service check pass/fail, duration, friction hit, final user TODO. Then one friction list deduplicated across both agents — friction only one agent hits is still friction, the weaker agent is the bar, fix the doc for both. Write the friction log to `$SPUR_SESSION_ARTIFACTS_DIR`.
SOURCE-INSTALL DEPLOY MODE
Gate: a change to `scripts/main-deploy.sh`, `deploy/*.service`, or the source-install deploy path ships only after this mode passes on the box. No merge without it.
Second use of the same box: run `scripts/main-deploy.sh` end to end against real systemd. The hermetic `tests/deploy` suite stubs `systemctl`, `curl`, `ss`, and `pnpm` — it proves nothing about real systemd, real startup timing, or a real Next build. Verified means the positive and the negative case below both pass. Planted agents play no part: skip steps 3-5, keep steps 2 and 6.
1 Reset in place (step 2). A box carrying the previous run's units and `.next` proves nothing.
2 Install node 22 and pnpm on the tested PATH. Create `/etc/spur/daemon.env` — main-deploy exits 1 without it. Passwordless sudo required; the script installs units with `sudo tee`.
3 `main-deploy.sh` resets its clone to `origin/main` every run, so `main` must BE the branch under test. Serve it from a local bare mirror:
git clone --bare https://github.com/<owner>/spur.git ~/spur-mirror
git -C ~/spur-mirror fetch origin '+refs/heads/<branch>:refs/heads/main'
git clone ~/spur-mirror ~/spur
cd ~/spur && pnpm main:deploy
4 Positive case: script exits 0, `spur-daemon.service` and `spur-web.service` both active, web answers 200 on 127.0.0.1:3012 (source-install port, not the npm path's 5555), and the run prints no `missing chunk` or `serving stale chunks` line.
5 Negative case: re-run and SIGTERM the script mid-build, then assert `spur-web.service` is still active — the exit trap restarts it. Two traps that read as a failure and are not:
- The stamp file matches after step 4, so a bare re-run takes the "Already deployed" fast path and never builds. Delete `~/.spur/main-deploy/repo/.git/main-deploy-last-successful` first, or there is no build to interrupt.
- `kill -TERM` on the `main-deploy.sh` bash pid does NOT fire the exit trap while `pnpm build` is the foreground child — bash runs the trap only after that child returns. Budget the remaining build (measured 630-690 s for a full one), never a 60 s poll. Confirm with `rc=143` plus the log's `main:deploy exiting rc=143 with spur-web inactive — starting`.
Then, and only then, check HTTP: `next start` needs 5-20 s after the trap starts the unit, so an immediate curl returns `000` on a healthy box. Assert the unit active first, poll `127.0.0.1:3012` for 200 second.
6 A guard keyed on the npm user units (`$HOME/.config/systemd/user/spur-daemon.service`) cannot fire on this box after step 2 — `tests/itest/reset-vm.sh:15` removes those units. Exercise it by planting the file by hand, then assert exit 1 and that both units' MainPIDs are unchanged.
AUTONOMOUS MODE
With no user present: run the full cycle above unattended for both agents, write the friction log to `$SPUR_SESSION_ARTIFACTS_DIR`, open a PR for any doc fix — never push straight to main — and report in-session. The box stays up after the run; never delete it, never stop it.
NOTES
The planted-agent + clueless-user role-play measures the docs, not your own knowledge — wanting to help the agent is a doc gap, write it down instead. Never hardcode the box's IP or the cloud project into this file — keep those in the local recipe, read it each run.
APPENDIX: BOOTSTRAP A REPLACEMENT BOX (ONE-TIME, USER INSTRUCTION ONLY, NOT PART OF THE CYCLE)
Only on the user's explicit instruction, only when the persistent box is gone or unrecoverable — needs cloud-CLI auth. Cheapest region near the user, e2-small floor, Ubuntu 24.04 LTS, key baked in (substitute your project/zone/key file):
gcloud compute instances create spur-itest \
--project=<gcp-project> --zone=<zone> --machine-type=e2-small \
--image-family=ubuntu-2404-lts-amd64 --image-project=ubuntu-os-cloud \
--boot-disk-size=30GB --boot-disk-type=pd-standard \
--metadata-from-file ssh-keys=<ssh-keys-metadata-file>
Fetch the external IP and record name + IP in the recipe. Confirm the default firewall exposes only SSH; app ports (4310/5555) must not answer from the public IP.
Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.
No comments yet. Be the first to comment!