Skip to content
Back to skills

Qa Verify Before Verdict

ASecurity

Use before any review_criterion approval or column move - the evidence each verdict needs, how to treat skipped, blocked and intermittent cases, and which exit a blocker takes

  • 109 stars
  • 0 votes
  • 0 copies
  • 0 views
  • Added September 19, 2026
ai-agentsrustgodockertestingapifrontendbackend

Works with

  • api

Security analysis

A100/100

Scanned October 4, 2026

npx -y skills add makifbaysal/tasktrooper --skill qa-verify-before-verdict --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Qa Verify Before Verdict?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Qa Verify Before Verdict
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/makifbaysal-qa-verify-before-verdict/badge)](https://www.skillsdirectory.com/skills/makifbaysal-qa-verify-before-verdict)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: qa-verify-before-verdict
category: qa
description: Use before any review_criterion approval or column move - the evidence each verdict needs, how to treat skipped, blocked and intermittent cases, and which exit a blocker takes
source: obra/superpowers (MIT), adapted
---
# Verify Before Verdict

No verdict without fresh executed evidence in this run.

- Every case on the task gets a result you recorded yourself: `passed` or `failed` with the exact command and its observed output (`actual`), UI cases additionally with a one-line evidence string (width/device, URL, viewport report, what it showed — never a screenshot path, there isn't one; frontend-manual-testing / mobile-manual-testing). The result belongs on the case (`set_test_case_result`), not in a comment — that is where the next reader looks, and it survives the run.
- A case you could not execute is `skipped` with what blocked it (no device, no stage deploy, missing credential), never `passed`. A case still `planned` when you try to hand the task on refuses the move.
- A pass verdict rests on ONE thing this iteration: every case executed manually by you, with its evidence. Writing automated tests is out of scope for now — do not add suites and do not hold a verdict waiting for one. The repository's existing pipeline is still read (`get_pipeline_status`): red is a finding, green is not a substitute for your own run.
- A criterion is approved (`approved=true`) only when you executed and observed it pass this run AND at least one `passed` test case carries its `criterion_id` — an approval with no linked passed case is what pm_uncovered_criterion rejects at PM UAT. A criterion you could not execute is never approved: say which one, what blocked it and what would unblock it.
- Both verdicts are moves out of `in_qa` — the column you took the task into before testing. All cases green → move it to pm_uat with the evidence on the cases and the criteria approved — except a task whose `task_type` field is `technical`: move it straight to human_uat instead, same evidence, same "no comment" rule. A backend-only, API-only or infra `task` or `bug` still goes to pm_uat; the field decides, never what the change touches. Any case fails → move it to need_revision with, per failure, the acceptance criterion, exact reproduction command, EXPECTED vs ACTUAL (and, for a visual defect, the width/URL and what the screen showed). A task left sitting in `in_qa` at the end of a run is an unfinished verdict, not a result.
- What actually fails a task vs. what is a note on a passing criterion: see bug-report-writing's rubric. Taste and polish beyond the team's UI floor are notes, never failures.

## Blocker triage

Not every "I could not verify this" exits the same way:

| Observation | Owner | Exit |
|---|---|---|
| Behaviour contradicts a criterion, human comment or the UI floor | dev | reject criterion, need_revision |
| Boot fails with the repo's own documented command (missing migration, crash on start, undocumented required env var) | dev | need_revision, quote command + log |
| Your environment: port busy, stale server, Docker absent, npm cache | you | fix it and re-run; never a finding |
| No device / no credential / no stage / ambiguous criterion | human | `ask_user` (after one `get_repo_tree` — it is refused until the run has read the repo), 1 call, choice mode with options such as "attach the device and resume", "provide a sandbox key", "accept reading A", "move to need_revision" |
| Device busy | system | stop; the task resumes by itself |

## Intermittent failures

A failure you saw once is not automatically a defect, and it is not automatically your environment either:

- Re-run it from a clean state (fresh page load / fresh request) up to 3 times.
- Fails ≥1 of 3 with the product's own behaviour → a real defect: report it "intermittent, 1/3", with the condition that seems to trigger it. Races are bugs, not noise.
- Fails only through your harness (your `sleep`, a selector that matched before render, a port collision) → fix the harness and re-run; never a finding.
- Wait for the condition, not a duration: `browser_wait_for` the post-action element, and for async poll with a timeout (`for i in $(seq 30); do <check> && break; sleep 1; done`). Never `networkidle` — Playwright's own docs mark it discouraged for assertions.

## Bug-type tasks

For `task_type=bug`, the reporter's own repro steps are case #1 and must pass on the branch — that's the red-green check that proves the fix actually touches the reported symptom, not a nearby path that happens to look related. When the pass looks ambiguous (intermittent, environment-dependent), confirm the same steps still reproduce the bug on the default branch or stage before trusting that the fix explains the difference.

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…