Audits PABX end-to-end tests (tests/e2e/) against the e2e-test-writer rules, producing a compliance report with a PASS/FAIL/PASS WITH WARNINGS verdict, per-violation severity, and fix suggestions. Read-only — never writes, fixes or runs tests. Dispatched by the evaluator or qa-preflight to confirm that an e2e test a writer corrected or newly wrote (a missing-coverage test, a guard) conforms before they trust it, or run directly by a user.
Scanned 9/30/2026
npx -y skills add dayvisonassis/sdd-skills --skill e2e-test-validator --agent claude-codeInstalls into .claude/skills of the current project.
Are you the author of E2e Test Validator?
Add the live security badge to your README — it updates automatically with every re-scan.
[](https://www.skillsdirectory.com/skills/dayvisonassis-e2e-test-validator)More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.
---
name: e2e-test-validator
description: Audits PABX end-to-end tests (tests/e2e/) against the e2e-test-writer rules, producing a compliance report with a PASS/FAIL/PASS WITH WARNINGS verdict, per-violation severity, and fix suggestions. Read-only — never writes, fixes or runs tests. Dispatched by the evaluator or qa-preflight to confirm that an e2e test a writer corrected or newly wrote (a missing-coverage test, a guard) conforms before they trust it, or run directly by a user.
---
# E2E Test Validator (PABX)
Audit **end-to-end test files** to ensure they comply with every rule the `e2e-test-writer`
enforces. **Read-only:** produce a compliance report; never modify test, harness or production
files, never execute tests (static analysis only).
> **Project scope:** PABX monorepo. The full rule set is in
> **`../e2e-test-writer/references/pabx-rules.md`** — validate against it.
**Scope:** `tests/e2e/`. Out of scope, reported as such: files under `tests/visual/`, and the
harness's own files as the e2e section of `GATES.md` names them — by default the seed tests
(`harness-seed.spec.js`), the runner config, the global setup and the shared fixtures module
(`tests/e2e/fixtures.js`), which must read stored sessions to build the API fixtures. A harness-like file shipped **with a feature's tests** (its own config,
global setup, or a support module shared between spec files) is not out of scope: it is a
CRITICAL H2 violation of the set.
## INPUT
- `test_file_path` (required) — test file or directory.
- `checklist_file_path` (optional) — the `e2e-test.md` for coverage cross-reference.
- `contract_path` (optional) — the feature's `contract.md`, to check that tags name real ids.
- `severity_filter` (optional) — `critical` | `major` | `minor` (default: all).
## OUTPUT
A **compliance report** (English): summary counts, overall verdict — `PASS` = 0 critical and 0
major (minors are listed, not counted against it); `PASS WITH WARNINGS` = 0 critical and ≥1
major; `FAIL` = ≥1 critical — each violation (`[SEVERITY] Rule-ID`, location, description,
expected, found, fix), traceability analysis, and positive findings. When dispatched, the
**verdict** is the signal the caller consumes.
---
## EXECUTION STEPS (3 Phases)
### Phase 1 — Structural Analysis
Inventory `describe`/`test`/hooks/fixtures (and their scope)/helpers and each test's tags. Map
checklist ↔ tests and contract ids ↔ tags (when provided). Identify every navigation, action,
assertion, request capture, `page.route` mock, API call (its URL base and whether its status is
checked), data-creation point and cleanup point. Note the file's path (which profile folder) and
its imports.
### Phase 2 — Rule-by-Rule Validation
Check EVERY rule from `../e2e-test-writer/references/pabx-rules.md`:
- **Forbidden Practices** [CRITICAL] — everything the writer's *Never* list forbids:
`waitForTimeout`, fixed sleeps or `networkidle` (V2); `retries` or `describe.configure({ retries })` (E4); a mocked happy path (V7); a filter validated only by an empty result (V5); a test with no action, load-only (V1); computed-style or box measurement (Boundary); creating or editing production code, the harness or `tests/visual/` (H2); direct database access — SQL, `knex`, `mysql2` (D2); `test.describe.serial`, state carried between tests, or a data-creating fixture that is not test-scoped (S5); acting on an item whose identity the test does not know — whatever renders first (`nth-child`, `.nth(n)`, `getByRole('option').first()`) — or XPath (V3). Choosing an item from data (a record the test created, an entry of a captured response) and acting on it by name or id is correct.
- **Traceability & placement** [CRITICAL] — the `evaluator` finds coverage only by tag and the session only by folder, so a miss here re-opens the gap on every evaluation: S1 the file sits in the profile folder whose session the flow needs, named per S1; S3 every test carries the anchored feature tag (the route's slug tag for work without a feature) and the id of every surface it touches, and a criterion tag only where it proves that criterion completely. With `contract_path`, every surface and criterion tag must name a real id.
- **Session** [CRITICAL] — A1 no login in a test, hook or fixture (no filling the sign-in form, no `POST` to the login route); A2 no forged token/session, no reCAPTCHA bypass; A3 no storage-state file read directly — another session only through its harness API fixture. The flow under test runs in the profile folder's session.
- **Data** [CRITICAL] — D1 no mutation of pre-existing customers, tickets, incidents or their relations, no child created under one (e.g. a ticket for an existing customer), and no submit pressed on a form bound to a pre-existing record — referencing pre-existing configuration (tenant, queues, categories, tags, custom fields) is allowed, editing it is not; D2 the test creates its own customer before its own ticket, with a run token, through the UI or HTTP API; D3 cleanup in the creating fixture's teardown (preferred) or `try/finally`, children before parents, UI-created records found through a parent the test owns, every removal attempted, each status checked, and a throw after the last attempt when any failed — collecting failures and throwing at the end is correct; a `catch` that swallows a failed removal is CRITICAL; D4 removal through a delete route and an API fixture; D5 no creation without a removal path; A3 API calls to the backend origin, never relative to the page, with setup statuses checked.
- **Missing data** [CRITICAL] — V4 no `test.skip`, early `return` or count-guarded branch that lets an absent precondition pass; the precondition is asserted.
- **Language & Naming** [CRITICAL] — S4 English in test/describe names, comments, variables and helper names; names state behavior. **Product copy is exempt:** Portuguese labels, button names and messages inside locators and assertions are correct, not a violation.
- **Checklist** [MAJOR] — when validating a feature's tests, its `docs/<feature-id>-<kebab>/e2e-test.md` exists, its coverage table lists every test title and records the D4 residue, and every row carries an outcome of the closed list in `../e2e-test-writer/references/pabx-rules.md` ("Coverage outcomes"), worded as that list words it — tests (`unproven` until proven), `disputed` entries naming one test each (alone or next to the row's tests), a `not e2e-testable` reason, `out of e2e scope — <suite>`, or `not in this request`. The `evaluator` reads coverage from it and rejects any other wording.
- **Structure** [MAJOR] — S2 CommonJS `require`.
- **Assertions** [MAJOR] — V3 role/label/text locators first; structural markers (component selectors, `name` attributes, component-scoped classes, rows filtered by visible text) only where the accessible name is unstable or absent. Requests captured before the action that fires them.
- **Requests** [MINOR] — V6 the requests the action should fire are captured and asserted; exactly-once asserted for every request whose duplicate would be a defect (a submit that creates, a search).
- **Execution** [MAJOR] — nothing that forces headed mode (E1).
V8 (never weaken an assertion) is a behavior of the writer, not a static property — flag it only
when a comment or diff context shows an expectation changed to match an observed defect.
Do not skip rules after finding criticals. State which were skipped and why.
### Phase 3 — Compliance Report
Emit the full report per OUTPUT. Verdict is derived strictly from the counts.
---
## RULES
**Always:**
- Validate against `../e2e-test-writer/references/pabx-rules.md`; check every rule.
- Give exact line/block locations and a concrete fix per violation; report positives. Output in English.
- Derive the verdict strictly from severity counts.
**Never:**
- Modify any test, harness or production file (read-only). Execute the tests. Mark PASS with any CRITICAL. Produce partial reports.
---
## Edge Cases
- **Empty test file:** CRITICAL. **Load-only test:** CRITICAL (V1).
- **Data created with no cleanup, or a cleanup that swallows a failed removal:** CRITICAL (D3). **Delete of an id the test did not create:** CRITICAL (D1).
- **Login in a `beforeAll`:** CRITICAL (A1) — one login per file still multiplies across files and spends the rate limit.
- **A Portuguese string inside `getByRole`/`getByLabel`/`toHaveText`:** product copy — not a Language violation.
- **File under `tests/visual/`, or a harness file of the repository:** out of scope — say so; do not validate it against these rules.
- **Directory input:** one report per file + a summary.
Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.
No comments yet. Be the first to comment!