Installs into .claude/skills of the current project.
Are you the author of Tests?
Add the live security badge to your README. It updates with every re-scan.
[](https://www.skillsdirectory.com/skills/anywh-sh-tests)
---
name: tests
description: Testing doctrine, test layout conventions, and where each kind of test lives in the anywh repo
---
# Testing Guidelines
## Testing Doctrine
Two kinds of test are preferred, in this order:
1. **Real integration tests** — real filesystem, real HTTP/WebSocket connections, real child processes where feasible. No mocks/stubs/fakes for anything the repo itself controls.
2. **Unit tests on pure/isolated logic** — pure functions or well-isolated modules where inputs and outputs are clear enough that no mock is needed.
Avoid mock-heavy tests that assert on implementation detail instead of behavior. If a piece of code can only be tested by mocking half its dependencies, that is a signal the code should be restructured to be more testable, not a signal to add more mocks.
Avoid tautological tests (tests that just restate the implementation). Favor tests that pin down an invariant or a boundary case that could plausibly break.
## The sanctioned mock boundaries: the `claude` process, `systemctl`, and the push gateway
The relay's job is to spawn `claude -p ...` per turn and stream its JSON events back over WebSocket. That external process is the one dependency that cannot be exercised for real in a test: it needs a live subscription, it is not deterministic, and it must never run in CI.
For relay integration tests, the sanctioned exception is a fake `claude` executable — a small script that emits canned JSON stream events on stdout — substituted for the real one via `PATH`/spawn override for the duration of the test. Everything else in that test (the HTTP server, the WebSocket layer, session persistence to disk, profile registry) runs for real, unmocked. This is the same shape of exception `mockAiRouter` plays in other coding-agent-desktop-app codebases: one deliberate seam at the LLM boundary, nowhere else.
A second boundary exists for the same reason, added after a real incident (2026-09-07): `DELETE /control/profiles/:id` (server.ts) shells out to the real `systemctl --user disable --now anywh-relay@<id>` to tear down a profile's systemd instance. Unlike every other side effect that route touches (the `.env` file, `profiles.json` — both already sandboxed per-test via `ANYWH_ENV_DIR`), `systemctl --user` talks to the operator's REAL systemd user session — there's no env var that scopes it to a test. A test exercising this route against the real binary disabled and stopped the operator's actual live `anywh-relay@trabalho` service and SIGKILLed an in-flight `claude` conversation running under it. `SYSTEMCTL_BIN` (mirroring `AGENT_BIN`'s existing pattern) now lets tests point this one spawn at `relay/tests/fixtures/fake-systemctl.mjs` (always exits 0, no real action) instead — `relay/tests/helpers/testServer.ts` sets it automatically for every integration test. **Never call `DELETE /control/profiles/:id` (or anything else that shells out to `systemctl --user` against real unit names) in a test without confirming `SYSTEMCTL_BIN`/`process.env.SYSTEMCTL_BIN` is pointed at the fake — the profile `id` used in a test can coincide with a real profile that exists on the machine running the test.**
A third boundary is the **push gateway** (`POST <gatewayUrl>`, contract C of `docs/push.md`), the service that turns the relay's notification into an APNs push. A real one needs a paid Apple account and would buzz a real phone, so the relay's own tests stand a local `node:http` server in for it: `relay/src/push/dispatcher.io.test.ts` drives the real `PushDispatcher` and the real `PushStore` (tmpdir) against it, with only the clock replaced so a backoff can be stepped through, and `relay/tests/push.test.ts` runs the whole path — a real relay and the fake `claude`, registration over HTTP, presence over the real WebSocket, a turn, then the request arriving at the fake gateway. The gateway fake only records what it was sent and answers what the test scripts (`200 {rejected, failed}`, `503`, `429`); nothing the relay itself controls is faked. The relay never talks to Apple, so there is no APNs fake here — that boundary belongs to the gateway's own repo.
## Where tests live
### `relay/`
`relay/src/**/*.test.ts` splits into two tiers by name, not by a different
practice — both run with `npm test` (`node:test`), both are unit-shaped,
the name is what tells a reader (or a future `import-x` rule) which
boundary a new test is allowed to touch:
- **`<module>.test.ts` — pure.** Nothing here touches `fs`, `spawn`,
`net`, or the clock (`Date.now()`/`new Date()`). If the module under test
needs one of those, either the test fakes it via a parameter/injected
function, or the test belongs in the next tier.
- **`<module>.io.test.ts` — boundary.** Touches a real tmpdir, a real
child process, or a real socket — but never a *new* mock. This repo has
exactly three sanctioned mock boundaries (below); a test that needs a fourth
one is a sign the module needs restructuring, not a sign to add a fake.
`sessionStore.test.ts` (real tmpdir) is this tier's shape today; new test
files that touch a boundary use the `.io.test.ts` suffix going forward,
and existing ones pick it up as they're next touched rather than in one
sweeping rename.
- `relay/tests/*.test.ts` — integration tests. Run with `npm run test:integration` (`npm run test:all` runs both). Each test file gets `server.ts` imported for its side effects against a throwaway `$HOME`/sessions file (`relay/tests/helpers/testServer.ts`, one call per file — `server.ts` has no exported bootstrap function and `node --test` already isolates each file in its own subprocess) and talks to it over a real WebSocket (`relay/tests/helpers/wsClient.ts`). `AGENT_BIN` points at `relay/tests/fixtures/fake-claude.mjs`, the fake executable described above — everything else is real.
- **A test that needs to observe the connection-time state burst (`cwd_state`, `permission_mode_state`, `draft_state`, `history_page`, `caught_up`, ...`SharedSession.addClient` sends) must attach its message collector in the SAME synchronous tick the socket is constructed, not after awaiting `connectSession`'s `open`.** Real finding writing the draft-persistence test: those frames are written synchronously inside the server's `connection` handler and often arrive in the same TCP read as the WS handshake response — `ws`'s client parser can emit several `message` events within the same call stack that fires `open`, before the `await` on a separate `connectSession()` call even returns control to the caller, silently dropping everything emitted in that window. Use `connectSessionAndCollectUntil` (wsClient.ts) for this, not `connectSession` + `collectUntil` as two separate steps — the plain two-step form is still correct for the (far more common) case of waiting for a response to something the test itself sends, since there the send always happens after the listener is already attached.
- **The embedded terminal (`/terminal`) talks to a real tmux via node-pty (terminalSession.ts) — no fake needed, tmux is deterministic and fast.** Two things that only show up testing it for real: (1) `spawnTerminal`'s `new-session -A ; set-option ... ; set-option ...` chain causes tmux to redraw the whole screen more than once right after attach — sending input before that settles can be swallowed, so wait for a quiet period (no new `data` frames for ~300ms) before typing, the same way a real user waits for the terminal to stop redrawing before they start typing. (2) a WS `close()` only detaches the tmux session (by design — that's what lets it survive a relay restart), it never kills it; a test that reconnects to prove persistence and then just closes the socket at the end leaves a REAL tmux server process running on the host. Always kill it via `POST /terminals/close` in a `finally`, unconditionally — a real finding while writing this test: three manual reruns without that left three live orphaned tmux servers on the sandbox running these tests.
- **Testing the "Stop" button (`stop_turn`) needs the fake claude to genuinely still be in flight, not racing to finish first.** The fixture supports this via `FAKE_CLAUDE_HANG=1`: after emitting the `system`/`init` event it blocks (no `assistant`/`result`) until it receives `SIGINT` — matches `runtimes/defs/claude/session.ts`'s tested claim that the real binary catches `SIGINT` and exits 0 with a valid `result` rather than dying raw. Two traps hit building this: (1) it's gated on `--output-format stream-json` specifically, not "any `-p` call" — title generation and next-message suggestion (`titleGenerator.ts`/`suggestionGenerator.ts`) fire their own `-p` calls in parallel with the same turn, with the identical prompt text but `--output-format text`, and nothing ever sends *them* a `SIGINT` — gating on the real turn's own output format only was what stopped those from hanging forever too. (2) the interrupted turn's synthetic `result` must carry `is_error: true` (not `false`) despite exiting 0 — `ClaudeSession.sendTurn` only classifies a turn as `stopped: true` via the error-result branch, so an `is_error: false` "successful" reply on SIGINT gets silently misreported as `stopped: false`, exactly backwards from a real interrupt.
### `client/`
- `client/src/**/*.test.tsx` — unit/component tests, colocated. Run with `npm test` (`vitest run`, `npm run test:watch` for the watch mode). `happy-dom` environment, `@testing-library/react` + `@testing-library/jest-dom` (config in `client/vitest.config.ts`, matcher setup in `client/tests/setup.ts`, ambient types re-exposed to `tsc` via `client/src/vitest.d.ts` since `tests/` sits outside `tsconfig.json`'s `include`). A component driven by a timer (`setTimeout`/`setInterval` inside `useEffect`) needs `vi.useFakeTimers()` plus the timer advance wrapped in `act()` from `@testing-library/react`, or the state update lands outside a render React Testing Library knows to flush.
- `client/tests/ui/` — integration tests that render the full `App` (via `client/tests/ui/helpers/renderApp.tsx`, which wraps it in `TooltipProvider` the same way `main.tsx` does — `<App />` alone throws, any Tooltip-using component needs the provider), drive it via click/type like a user, and mock only the network edge (`client/tests/ui/helpers/fakeRelay.ts` stubs `WebSocket`/`fetch` — the client-side equivalent of the relay's `claude`-process boundary). Never assert on or reach into React state/context directly to shortcut a test; go through the UI. Known traps hit building this tier, now handled once in `client/tests/setup.ts` (or the fake relay helper) instead of per-test:
- **Virtualized lists render zero rows in happy-dom.** `MessageLog.tsx` uses `@tanstack/react-virtual`, which measures the scroll container via `offsetWidth`/`offsetHeight` (its own `getRect`, not `getBoundingClientRect`) — happy-dom has no real layout engine, so both are always `0`, and the virtualizer computes an empty visible range regardless of how many entries exist in state. Fixed in `tests/setup.ts` by stubbing `HTMLElement.prototype.offsetWidth`/`offsetHeight` (plus `getBoundingClientRect` and a `ResizeObserver` that fires once synchronously — needed for other measurement paths even though the *initial* read turned out to be the actual blocker here). If a future virtualized component still renders empty under this tier, suspect a measurement this stub doesn't cover before suspecting the test's data.
- **`caught_up` is a hard gate, not an optional nicety.** `useRelayClient`'s `ready` flag only flips true on `{"type":"caught_up"}` — until then `ChatPanel` shows a skeleton and won't render even the user's own just-sent bubble. Any fake relay for this tier must send it after `open`.
- Tiptap's composer is a bare `[contenteditable]` with no explicit `role="textbox"` — query it with `findByLabelText`, not `findByRole("textbox", ...)`.
- Radix's portal-based components (`Dialog`, `Popover`, `Tooltip`, `DropdownMenu`) render into `document.body`, not the app's root container — a query scoped to the rendered root will miss them; query `document.body` (or use `within(document.body)`) for portal content.
- `client/tests/e2e/*.spec.js` — end-to-end tests via WebdriverIO + `@wdio/tauri-service` (embedded provider), driving the real Tauri app (real window, real webview) on Windows/Linux/macOS with no external driver needed on any of the three. This is the only tier that can catch Tauri-webview-specific bugs (e.g. `window.confirm` unreliability, Radix focus-trap breaking under WKWebView — both already burned this project once, see `CLAUDE.md`). Expensive to run (builds the whole app), so CI only runs it when a PR/push actually touches `client/**` (see "CI" below) — not unconditionally on every PR. It should also gate on release builds once that workflow exists (not built yet).
- Requires a special build first: `npm run build:e2e` (`tauri build --debug --no-bundle --features e2e --config src-tauri/tauri.e2e.conf.json`), then `npm run test:e2e` (`wdio run ./wdio.conf.js`). On Linux this needs a display — `xvfb-run -a npm run test:e2e` (Xvfb already assumed available; it's what CI runners need too).
- The `e2e` Cargo feature (`client/src-tauri/Cargo.toml`) gates `tauri-plugin-wdio` + `tauri-plugin-wdio-webdriver` — both `optional = true`, so neither is even fetched/compiled for a normal `tauri dev`/`tauri build`, not just inert. `tauri.e2e.conf.json` is a `--config` override (never touches the real `tauri.conf.json`) that adds the one extra capability (`wdio:default`) those plugins' frontend-invokable commands need — `tauri-plugin-wdio-webdriver` itself needs none (empty permission set; it drives the app over the WebDriver protocol directly, not through Tauri's IPC/capabilities layer).
- Validated end-to-end on Linux (Debian, WebKitGTK, under Xvfb): first on 2026-09-07 with a relay-independent smoke test (app launches, sidebar renders), and again on 2026-09-12 with the shell suite green in ~5s for the whole run. Not yet validated on Windows/macOS.
- **This tier has three halves that must all be present, not two**: the `tauri-plugin-wdio` crate (Rust, behind the `e2e` feature), `@wdio/tauri-service` (test side), and `@wdio/tauri-plugin` (frontend, a `devDependency`), whose only job is to snapshot `window.__TAURI__.core` into `window.__wdio_original_core__` before anything can Proxy it. The frontend half was missing from 2026-09-07 to 2026-09-12, and it was never the harmless curiosity this file used to claim it was: the service's per-command focus probe (`ensureActiveWindowFocus` → `getWindowStates`) invokes `plugin:wdio|get_window_states` through exactly that global, so **every element lookup waited out a 5s timeout (~10s nested) and then gave up**, logging `Tauri core.invoke not available after 5s timeout`. Only the `focusCommands` list pays it (`getTitle`, `findElement(s)`, `$`, `$$`, `elementClick`) — `browser.execute` does not, which is why ad-hoc debugging felt fast while the specs crawled. Two lookups still fit inside mocha's 60s budget, so the smoke test kept passing while any richer spec died of a bare `Error: Timeout` naming no element. Wiring the frontend half took the whole suite from 1m54s to ~5s. **If that warning ever reappears in a run, the tier is silently ~10s/command again — treat it as a build regression, not as noise.**
- The frontend half is gated on a Vite *mode*, not a shell env var, so it behaves the same on Windows and macOS: `tauri.e2e.conf.json` overrides `beforeBuildCommand` to `npm run build:frontend:e2e` (`vite build --mode e2e`), which loads `client/.env.e2e` (`VITE_E2E=true`), which flips the one `import.meta.env.VITE_E2E` branch in `client/src/main.tsx`. For a normal dev/build that branch is statically false and Vite drops it with its dynamic import — verified by grepping a production bundle for `wdio` (no hits). A new key there also needs an entry in `client/src/vite-env.d.ts`, whose `ImportMetaEnv` replaces Vite's default (no index signature), or `tsc` rejects the read.
- **A WebDriver click in this webview cannot open a Radix dropdown, and this is a driver limit, not an app bug.** Recording native listeners during a real `element.click()` shows it delivers only `click` (with `isTrusted: false`) plus a trusted `focus` — no `pointerdown`, no `mousedown`, even though the webview does implement `PointerEvent`. Radix's `DropdownMenuTrigger` opens on `pointerdown`, so it never fires; the W3C Actions API (`browser.performActions`) does not help either. The visible symptom is deceptive: the trigger's *tooltip* opens on that focus event (`data-state="instant-open"`, which in Radix means focus-opened rather than hover-opened), so it looks like the menu is broken or the tooltip is stealing the click. Drive such triggers with `browser.keys("Enter")` after focusing, which still exercises the portal and focus trap this tier exists to test. Before concluding a control is broken here, record the events — do not infer from the DOM alone.
- **The driver cannot type into a `contenteditable` either, and this is the same class of limit.** With the editor verifiably focused (`document.activeElement` is the `.ProseMirror` div and carries `ProseMirror-focused`), all three ways in leave `textContent` empty: `browser.keys`, `element.addValue`, and a `document.execCommand("insertText")` run inside `browser.execute`. Key *events* do arrive — that is what `browser.keys("Enter")` relies on to open a Radix menu — but WebKit's editing pipeline never runs for synthesized ones, so no `beforeinput` reaches ProseMirror and no text is inserted. Any flow that needs typed text (sending a message, the composer's slash menu, the typo confirmation) belongs in `client/tests/ui`, not here. Related: element focus is not assertable across commands either — read back inside the same `browser.execute` that opened a conversation, the editor is focused; read from any later command, it never is.
- WebdriverIO's bare `*=text` partial-text selector silently resolves to nothing under this driver (returns false in ~17ms, no wait involved) even when the text is in the DOM. A tag-qualified `span*=text` works, and a tag-agnostic `//*[contains(text(), "text")]` XPath works and is what the specs use — it does not pin the assertion to whatever element the text renders as today.
- Running it from a fresh git worktree needs two gitignored build inputs the checkout will not have: `client/node_modules` and `client/src-tauri/binaries/` (the `tailnet-sidecar` binaries, or the build fails with ``resource path `binaries/tailnet-sidecar-<triple>` doesn't exist``). Note the `.gitignore` entry for that directory carries a trailing slash, so it matches directories only — satisfying it with a *symlink* makes it trackable and `git add -A` will happily commit an absolute local path. Copy the host triple's binary into a real directory instead. Also export `PATH=$HOME/.cargo/bin:$PATH` for any non-login shell (a background job runner, CI step), or `tauri build` fails early with `failed to run 'cargo metadata'`.
- No relay is running during these tests today — flows that need one (send a message, etc.) aren't covered by this tier yet; it currently only proves the app itself boots and its shell renders under the real webview.
## When to test
New features and bug fixes should ship with tests. When fixing a bug, write the failing test first (reproduce the bug), then fix the code, then watch the test pass.
There is no coverage percentage target. The goal is covering main flows and the invariants/boundary cases that could actually break, not maximizing a number. A simplification that removes code usually does not need a new test; added complexity usually does.
## CI
`.github/workflows/pr.yml` runs on every push to `main` and every PR. A `changes` job (`dorny/paths-filter`) decides what else runs: `relay/**` touched → the relay job (`npm run build` + `npm run test:all`); `client/**` touched → both `client-unit` (`npm run build` + `npm test`) and `client-e2e` (installs the Linux Tauri/audio/whisper system packages, `npm run build:e2e`, then `xvfb-run -a npm run test:e2e`). Only Linux runs in CI today (matching the only platform this tier has actually been validated on, 2026-09-07) — extending `client-e2e` to a Windows/macOS matrix is future work, not done yet.
## Determinism
Tests run in varied environments (including slower CI runners). Prefer explicit synchronization (wait-for-condition helpers) over arbitrary `sleep`/timeout-based waits, which are the main source of flaky tests.