Skills DirectorySkills Directory
SkillsLearnSecurityCategoriesDocsCommunityBlog
Sign InSubmit Skill
Skills Directory

Security-tested agent skills for Claude, coding agents, and AI workflows.

Directory

  • Browse Skills
  • All Skills A–Z
  • Claude Skills
  • Claude Code Skills
  • Agent Skills
  • Categories
  • Authors
  • Submit a Skill

Learn

  • Learn Hub
  • Install Claude Skills
  • Write SKILL.md
  • Skills vs MCP
  • Directories Compared

Security

  • Security
  • Methodology
  • Secure Claude Skills
  • Security Badges

Company

  • About
  • Community
  • Blog
  • API Docs
  • Advertise

2026 Skills Directory. All rights reserved.

ProTermsPrivacyRefunds
Back to skills

Writing E2e Tests

ASecurity

Canonical rules for writing, rewriting, or reviewing Shift desktop Playwright E2E tests and visual goldens under `apps/desktop/e2e/`. Use whenever you add or change a `.spec.ts`, a fixture, `EditorDriver`, a screenshot baseline, or Playwright project membership, and whenever you investigate a flaky or failing E2E test. Covers waits, oracles, goldens, fixtures, projects, and flake verification.

274 stars
0 votes
0 copies
0 views
Added 9/27/2026
developmentgonodeperformance

Works with

cli

Security Analysis

A100/100

Scanned 9/27/2026

Install to Claude Code

$npx -y skills add shift-editor/shift --skill writing-e2e-tests --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Writing E2e Tests?

Add the live security badge to your README — it updates automatically with every re-scan.

Security grade badge for Writing E2e Tests
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/shift-editor-writing-e2e-tests/badge)](https://www.skillsdirectory.com/skills/shift-editor-writing-e2e-tests)

More formats (shields.io, HTML) on the badges page.

Download with Pro
Files
SKILL.md
---
name: writing-e2e-tests
description: Canonical rules for writing, rewriting, or reviewing Shift desktop Playwright E2E tests and visual goldens under `apps/desktop/e2e/`. Use whenever you add or change a `.spec.ts`, a fixture, `EditorDriver`, a screenshot baseline, or Playwright project membership, and whenever you investigate a flaky or failing E2E test. Covers waits, oracles, goldens, fixtures, projects, and flake verification.
---

# /writing-e2e-tests — How desktop E2E tests are written

An E2E test is worth its cost only when it fails for exactly one reason: the user-visible behavior it names is broken. It must not fail because the machine was slow, the theme changed, a sidebar moved, or a retry happened to pass.

Read `/writing-tests` first for the general rules (observable state, no mocks, migration ledgers). Read `apps/desktop/e2e/README.md` for the full fixture and project reference. This skill is the checklist that turns those into a test that stays green for the right reasons.

## Choose the layer before writing anything

| The behavior is…                                                                                          | Owning test                                       |
| --------------------------------------------------------------------------------------------------------- | ------------------------------------------------- |
| Pure computation (layout math, latch rules, range selection)                                              | Pure unit test — extract the function first       |
| Observable through editor state after a click, drag, or key                                               | `TestEditor` tool or command test                 |
| A real browser or Electron contract (DOM events, focus, native menus, dialogs, windows, processes, files) | Semantic Playwright test                          |
| Appearance: colours, strokes, handles, layout                                                             | One focused golden, after a semantic precondition |
| Hardware rendering, GPU residency, WebGPU presentation                                                    | `gpu` project test                                |
| Save, quit, crash, recovery, file activation, cross-OS paths                                              | `platform` project test                           |

If most assertions in a draft E2E test read editor state after `page.evaluate`, the behavior probably belongs in `TestEditor`. Keep the E2E test as a thin proof that the real UI reaches it.

## Waiting: never sleep

`waitForTimeout` is banned. Wait for the condition that makes the next step valid:

| Before you…                                                                                      | Wait for                                                                                                           |
| ------------------------------------------------------------------------------------------------ | ------------------------------------------------------------------------------------------------------------------ |
| Read geometry, project coordinates, or send pointer input after setup changed geometry or camera | `editor.waitForCanvasRender()`                                                                                     |
| Assert after raw `page.mouse` moves                                                              | `editor.flushPointerMoves()`, then `waitForCanvasRender()` or `waitForIdle()`                                      |
| Act on an editor route                                                                           | `waitForEditorReady()` / `openCatalogGlyph()` (works for authored and preview sessions)                            |
| Click a catalog cell by coordinates                                                              | `clickFirstCatalogGlyph()` — it waits for a laid-out, settled Grid                                                 |
| Act on a workspace                                                                               | `waitForWorkspaceReady()`                                                                                          |
| Assert after quit or close                                                                       | `dirtyDocumentDecisions()` / `dirtyDocumentRequests()` — main-process records, not elapsed time                    |
| Compare Grid frames after an edit                                                                | Published Grid state (`data-grid-readiness`, `data-atlas-build-count`, active location), armed after setup settles |

When elapsed time is itself the contract (debounce, momentum, idle timeouts), put the rule in a pure function driven by timestamps and unit-test it. If the E2E test must remain, dispatch events inside the page and wait on the page clock (`performance.now()`) that stamps them. `glyph-view.spec.ts` zoom momentum is the template.

Negative assertions ("nothing else happened") need a positive anchor: wait for the event that would have triggered the unwanted outcome, then assert the outcome is absent.

## Oracles: assert what the editor decided, not what pixels look like

Never count or sample hard-coded colours from a canvas. Colour oracles break under themes, alpha, anti-aliasing, device scale, and renderer changes, and many pass when the feature is broken (an empty canvas has zero "wrong" pixels).

Use published state instead:

- geometry: `editor.outline()`, `pointPosition()`, `selectionIds()`, `selectionBounds()`;
- interaction: `toolState()`, `hoverId()`, `activeSnapGuides()`;
- rendering decisions: `visibleOutlines(nodeId)` and `handleStates(node)` on the glyph node definition;
- UI state: roles, accessible names, `aria-pressed`, `aria-checked`, `aria-disabled`, `data-*` state published for tests;
- persistence: `savedGlyphNames()` for `.shift` files and `exportedGlyphNames()` for exported fonts — never `existsSync`, byte inequality, or `size > 0`.

If the state you need is not published, publish it from the production code that renders it (same computation, not a parallel reimplementation), add a unit test for it, then assert it. That is how `visibleOutlines` and `handleStates` were added.

Assertions must be exact enough to fail on the wrong answer:

- assert the exact set or order, not `count() > N` or `not.toEqual(before)`;
- derive expectations from the fixture or the UI (for example, read the ordered category buttons and compute the expected range) instead of hard-coding one neighbour;
- run the mental-deletion check: if the feature were a no-op, would this assertion fail?

## Goldens: one contract, one focused capture

All goldens go through `fixtures/snapshots.ts`; `scripts/check-e2e-projects.mjs` rejects `toHaveScreenshot` or `toMatchSnapshot` anywhere else.

- `expectCanvasSnapshot(editor, name)` — the editor canvas stack, exact comparison, after the canvas has rendered.
- `expectPanelSnapshot(locator, name)` — one panel, menu, toolbar, or dialog.
- `expectPageSnapshot(page, name)` — full window; use sparingly, only for overall chrome.

Before every golden:

1. Assert the semantic precondition that makes the image meaningful: the point count, selection, tool state, outline targets, handle states, zoom, active theme (`data-color-theme`), or `devicePixelRatio`.
2. Park the pointer off the captured element unless hover is the contract.
3. Capture the smallest element that shows the contract. Do not use a full-page golden to protect a toolbar button.

Rules:

- One golden per distinct visual contract. Do not add near-duplicate images; if two captures are byte-identical, one of them is redundant.
- Canvas goldens use `maxDiffPixels: 0` with `threshold: 0.02` on every host: enough for antialiasing differences between hosts (≤ 0.009), far below a real colour-token change (~0.07). Never raise either to make a mismatch pass; the default threshold of 0.2 silently accepts token changes.
- Interface goldens (`expectPanelSnapshot`, `expectPageSnapshot`) are exact but compared on CI only, with baselines generated on the runner via the `ci: update visual snapshots` label. Never commit a locally generated interface baseline. Prefer semantic assertions and keep interface goldens few.
- Goldens never pass on retry. The helpers throw on a retry attempt, so explain the first failure.
- Use canvas-local positions (`dragCanvas()`, `canvasPagePoint()`) so layout changes cannot redraw geometry.
- Theme goldens select the theme through the product and assert `data-color-theme` first. HiDPI goldens use `test.use({ deviceScaleFactor: 2 })` and assert `devicePixelRatio` first. Add one golden per palette branch or rendering path, not one per theme.

Updating baselines: only for an intentional appearance change. Run the focused spec with `--update-snapshots -g "<title>"`, open every changed PNG and describe what changed, then rerun without update mode and with `--retries=0 --repeat-each=3`. Canvas baselines may be generated locally. Interface baselines are generated only by the `ci: update visual snapshots` label; inspect the runner-generated images before merging.

## Fixtures and isolation

- Use the shared `editor` fixture; construct `EditorDriver` directly only for additional windows.
- Relaunch through the `relaunch` fixture, never raw `electron.launch`. The process registry owns every launched tree, attaches diagnostics, and kills it even when the test times out.
- Every launch environment comes from `shiftTestEnvironment()`. Never spread `process.env` into a launch.
- Script native dialogs with the fixture options (`dirtyDocumentChoice(s)`, `saveShiftPaths`, `openFontPath`). An unexpected renderer dialog fails the test; opt in with `allowRendererDialogs` only when the dialog is the behavior.
- Put shared setup in fixtures or `EditorDriver` (`openScratchGlyph()`, `commitInputValue()`, `toolButton()`), not copied helpers in specs. Do not grow `EditorDriver` with one-off helpers.
- Prefer `getByRole()` / `getByLabel()` and locators from `fixtures/appLocators.ts`. No parent traversal (`locator("..")`), no styling classes, no `toHaveCSS()` unless the style value is the product contract.

## Structure

- One behavior per test. Split long flows when an early failure would hide unrelated coverage.
- Do not accumulate state across unrelated steps; start each test from a fixture.
- Setup may use `page.evaluate` or `insertContent`; the behavior under test must go through the user surface.
- Name tests after the behavior: "shift-selects glyph categories as a visible range", not "category test".

## Projects

Membership is explicit in `apps/desktop/playwright.config.ts`. Add a new spec to the right list — `visual`, `platform`, `gpu`, or `perf` — and run `node scripts/check-e2e-projects.mjs`. A spec that needs native lifecycle behavior on Windows and Linux belongs in `platform`. A spec whose only GPU dependency is incidental belongs in `visual`.

## Proving a test is not flaky

Before committing a new or changed E2E test:

```sh
pnpm test:e2e:visual e2e/<spec>.spec.ts -g "<title>" --repeat-each=10 --retries=0
```

For timing-sensitive flows, repeat under CPU load (for example, several `yes > /dev/null &` processes, killed afterwards), because CI runners are slower than development machines. For platform behavior, rely on the Windows and Linux CI jobs and say so in the pull request.

When a test fails or flakes:

1. Read the trace and `electron-diagnostics` before changing anything.
2. Classify it: product regression, test oracle, or infrastructure. Do not call a single failure flaky without a retry pass or a rerun on the same commit.
3. Fix the oracle or the wait. Never add a sleep, a retry, or tolerance, and never regenerate a baseline without understanding the diff.

The CI "E2E Report" job lists tests that passed only on retry. Treat every entry as a bug to fix, not noise.

## Review checklist

- [ ] The layer is right; logic that `TestEditor` can observe is tested there.
- [ ] No `waitForTimeout`, no fixed delays, and no polling without a semantic condition.
- [ ] No colour counting, pixel sampling, byte inequality, or "something changed" oracle.
- [ ] Assertions are exact and fail when the feature is a no-op.
- [ ] Every golden goes through `fixtures/snapshots.ts`, captures one contract at the smallest element, and follows a semantic precondition.
- [ ] Changed baselines were inspected and verified with `--retries=0`.
- [ ] Launches, relaunches, and dialogs go through fixtures.
- [ ] The spec belongs to the intended project, and the project check passes.
- [ ] The test passed `--repeat-each=10 --retries=0`, and the pull request lists the exact E2E commands and any coverage not run.

Attribution

shift-editorshift-editor
View sourceMore from shift-editor →
SSkills DirectorySkills Directory

Know which skills are safe — weekly.

Best new skills + every skill we flagged as malicious. From the team that scanned 103,619.

Join free

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments (0)

No comments yet. Be the first to comment!

SSkills DirectorySkills Directory

Know which skills are safe — weekly.

Best new skills + every skill we flagged as malicious. From the team that scanned 103,619.

Join free

Related Skills

Browser Extension Developer

Use this skill when developing or maintaining browser extension code in the `browser/` directory, including Chrome/Firefox/Edge compatibility, content scripts, background scripts, or i18n updates.

284972 votes

Seo Optimizer

SEO optimization with keyword analysis, readability assessment, technical validation, content quality. Use for search rankings, blog posts, content audits, or encountering keyword density, readability scores, meta tags, schema markup errors.

2222 votes

Google Official Seo Guide

Official Google SEO guide covering search optimization, best practices, Search Console, crawling, indexing, and improving website search visibility based on official Google documentation

1862 votes

Tanstack Start

Build a full-stack TanStack Start app on Cloudflare Workers from scratch — SSR, file-based routing, server functions, D1+Drizzle, better-auth, Tailwind v4+shadcn/ui. Use whenever the user mentions TanStack Start, asks to scaffold a full-stack Cloudflare app with SSR, wants an SSR dashboard, or asks for a React 19 + Cloudflare Workers app with file-based routing and server functions — even if they don't name TanStack Start specifically. No template repo — Claude generates every file fresh per ...

10311 votes

Pentest

PTES-aligned adversarial security audit for backend, frontend, and mobile applications. Produces a CVSS-scored Hacker Report with verified PoCs and phased remediation.

5491 votes
View all in development →