Skills DirectorySkills Directory
SkillsLearnSecurityCategoriesDocsBlogPro
Sign InSubmit Skill
Skills Directory

Security-tested agent skills for Claude, coding agents, and AI workflows.

Directory

  • Browse Skills
  • All Skills A–Z
  • Claude Skills
  • Claude Code Skills
  • Agent Skills
  • Categories
  • Authors
  • Submit a Skill

Learn

  • Learn Hub
  • Install Claude Skills
  • Write SKILL.md
  • Skills vs MCP
  • Directories Compared

Security

  • Security
  • Methodology
  • Secure Claude Skills
  • Security Badges
  • Chrome Extension
  • Skill Manager

Company

  • About
  • Community
  • Blog
  • API Docs
  • Advertise

2026 Skills Directory. All rights reserved.

ProTermsPrivacyRefunds
Back to skills

Sumo Qa Implementing With Tdd

ASecurity

Use after sumo-qa-deciding-approach picks tdd-scaffold, regression-first, or coverage-first-then-refactor — e.g. "write a regression test for this bug" or "scaffold the failing tests first". Walks plan → name-the-risk-and-test-idea → confirm → red → hand off → green → review, one section per turn with confirmation gates. Don't write the test until the test idea has been agreed.

7 stars
0 votes
0 copies
0 views
Added 9/23/2026
testingpythongoshelltestinggitapi

Works with

cliapi

Security Analysis

A100/100

Scanned 10/4/2026

$npx -y skills add sumithr/sumo-qa --skill sumo-qa-implementing-with-tdd --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Sumo Qa Implementing With Tdd?

Add the live security badge to your README — it updates automatically with every re-scan.

Security grade badge for Sumo Qa Implementing With Tdd
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/sumithr-sumo-qa-implementing-with-tdd/badge)](https://www.skillsdirectory.com/skills/sumithr-sumo-qa-implementing-with-tdd)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
Files
SKILL.md
---
name: sumo-qa-implementing-with-tdd
description: Use after sumo-qa-deciding-approach picks tdd-scaffold, regression-first, or coverage-first-then-refactor — e.g. "write a regression test for this bug" or "scaffold the failing tests first". Walks plan → name-the-risk-and-test-idea → confirm → red → hand off → green → review, one section per turn with confirmation gates. Don't write the test until the test idea has been agreed.
---

# Implementing with TDD

The user has product context (what "wrong" looks like, the API shape) the AI can't infer from code: surface it through questions, don't assume.

## Output discipline (mandatory)

Inherits the global discipline from `using-sumo-qa`: **output discipline** (never surface internal taxonomy labels — say *"behaviour change in pricing"*, not *"Classification: business_logic_change"*), **output economy** (spend output on findings not framing; no preamble or self-narration; one question per turn; no closing pleasantries), knowledge authority hierarchy, internal scaffolding stays internal, and specialty-tool fit.

<HARD-GATE>
Do NOT write the failing test in the same turn you propose the test idea. Walk: risk → assertion shape → smallest failing test → confirm → write → run → show red. Tests written before the user agrees what they catch are guesses, not red-phase evidence.
</HARD-GATE>

## The Iron Law

**RED PHASE FIRST. NO PRODUCTION CODE BEFORE A FAILING TEST.** A test that has never failed has never tested anything — the red phase is the proof.

**Stub allowance — narrow.** A production-side stub is permitted in the red phase ONLY when the test can't otherwise be collected (e.g. the function under test doesn't exist yet, so the test file fails at import). It must be signature-only: `def apply_discounts(order): raise NotImplementedError` or `: pass`. **Any behaviour in the stub — a partial implementation, a heuristic return, a branch that happens to satisfy the assertion — is an Iron Law violation:** the red phase would then prove the stub matches the assertion, not that the test catches the bug.

## When to Use

`sumo-qa-deciding-approach` routes here for: `tdd-scaffold` (greenfield-ish new behaviour), `regression-first` (bug fix — reproduce as a failing test first), or `coverage-first-then-refactor` (behaviour-preserving refactor — characterization tests pin behaviour BEFORE the refactor).

For `strengthen-test-coverage` (mutation follow-up), route to `sumo-qa-strengthening-tests` instead — different discipline (production code stays locked).

## Checklist

You MUST work through these in order. Steps 1–3 are AI-only homework; the user's confirmation gates step 4 onward.

1. **Name the risk this cycle targets** *(no user question)* — show the user that named risk in plain words; the approach stays internal, never a label naming it. If no risk was named, route to `sumo-qa-preparing-for-work` first.

2. **Walk the repo for the target** *(no user question)* — use the host's file tools. Find (a) the production file touched, (b) the matching test file (or where one belongs), (c) the existing test style (framework, fixtures, assertions), (d) for regression-first: the failing path that reproduces the bug. Don't ask "what test framework?" — read a sibling test file. For `coverage-first-then-refactor`, if the repo emitted a coverage artifact (any format — `coverage.json`, Cobertura, `lcov.info`), read it with the host's file tools to LOCATE the uncovered CHANGED lines/functions worth characterizing before the refactor; the artifact only points at candidates — coverage % never decides what matters (step 3).

3. **Pick the smallest failing test idea** *(no user question)* — first call `sumo_qa_load_techniques()` if not already loaded (this skill owns that load; the router does not pre-load it). Then name (a) the target test file path, (b) the function under test, (c) the input that triggers the risk, (d) the assertion that distinguishes broken from fixed, and (e) the test-design technique from that loaded `sumo_qa_load_techniques()` catalogue, justified by this risk. The technique MUST be the verbatim catalogue heading (lowercase, with the catalogue's suffix — e.g. "decision tables", "state transition testing", "boundary value analysis", "equivalence partitioning", "exploratory testing", "pairwise testing"); don't paraphrase, title-case, or shorten it. **Tautology check:** if the assertion re-states the production code, pick an observable outcome instead. **Setup-discriminator check:** if the test setup makes broken and fixed implementations produce the same outcome, redesign it so they diverge — e.g. a mock that rejects on first call but resolves on second when testing rejection-cache invalidation. **Expected-value derivation check:** derive the expected value from the SPECIFIC input by applying the domain rule, never recall a generic fact. Trace input → rule → expected in one line before pinning (e.g. *"Jan 31, anchor 31 → Feb 2024 has 29 days (leap year) → expected 2024-02-29"*). If the trace doesn't name the input-specific fact (leap-year status for THAT year, length of THAT month, offset of THAT date, locale of THAT string), the expected value is a guess — a broken impl that hardcodes the same generic (Feb=28, all-months=30, UTC=+0) passes the assertion and the test never discriminates.

   - For `coverage-first-then-refactor` characterization tests, every fixture value (strings, numbers, identifiers) MUST be copied verbatim from the ground-truth context — never paraphrase or shorten; the asserted output must be exactly what the function currently produces.
   - For characterization tests, prefer techniques that pin existing behaviour: `equivalence partitioning` or `exploratory testing` charters. Avoid `use case testing` — that fits new-behaviour scaffolding, not pinning existing behaviour.
   - When the function under test parses an external CLI/API's output (greps stdout, regex-matches a body), the technique is `real-capture fixtures for external-output matchers`: capture the tool's REAL output to the fixture FIRST, THEN write the matcher. An invented fixture validates your assumption, not the real contract.
   - When the risk is "the BUILT artifact ships (or omits) a required member" (a wheel's `py.typed`, an image's stripped toolchain), the technique is `build artifact contents verification`: build and open it, then assert membership against the BUILT output, never the source tree.
   - When code spawns a tool that can inherit redirecting vars (e.g. git: an absolute `GIT_DIR`/`GIT_INDEX_FILE` from a hook or git wrapper; uv: `VIRTUAL_ENV`, `UV_*`; pip: `PIP_*`; bare `pip`/`python`: `PATH`), `hermetic subprocess environment` is the technique if inherited env is the defect, else a risk beside the primary. Input: them set to a throwaway value; fix: an explicitly constructed env.

4. **Confirm the test idea, only for the AMBIGUOUS parts** — name target, fixture style, and proposed assertion, then ask ONE focused question for what code couldn't answer (e.g. *"is 90.0 right, or does VIP stack with promo?"*). If unambiguous, skip the question.

5. **Write the failing test** — use the host's edit tool. Do NOT ask the user to write it. Match the sibling tests' framework and fixture style. Name its step-3 technique with it.

6. **Run the test and SHOW THE RED OUTPUT** — from a real run only, capture the actual assertion failure (expected vs. got, line number). Import/syntax/fixture errors are NOT red — adjust until you see a real assertion failure for the right reason. Surface verbatim. **No command tool, or your run attempt was denied by the user or host policy?** Don't delegate the run, ask the user to run it, trace the failure by hand, or write production code: go to step 7's second phrasing. A failed command isn't that: fix it, re-run.

7. **Hand off to the user, then stop** (retrospective regression: see its item 4) — the LAST line of your reply is EXACTLY one of these two phrasings, with nothing after it (no pleasantries, no confirmation question, no "shall I"):
   - if step 6 really ran the test and you pasted its assertion failure: "red phase confirmed. Implement to make it green; I'll re-run when ready. If you'd like me to write the production code, say so."
   - otherwise: "I'll run this and surface the assertion failure next."
   A real run is a command-tool call you made this turn. Without one, write no red-output block, no production code and no commands for the user to run: the second phrasing is the whole handoff.
   No run means no failure text anywhere, prose included: never write what a test runner would print (`AssertionError: ...`, a traceback, a `FAILED` line, a line number) or an "Expected red output" section. That text is a run's evidence; writing it yourself forges it and hides the one thing only a run shows: whether the test fails for the right reason. The step-3 trace justifies the expected value, never a predicted failure. To say why the test discriminates, state in words what broken and fixed code return for this input.
   - Bad: *"the buggy code returns None, so it fails: `AssertionError: assert None == Path(...)` at line 24."*
   - Good: *"For this input the buggy code returns None; the fixed code returns the toolkit path."*, then the second phrasing.

8. **Re-run after green-making change** — confirm it passes for the right reason (not a weakened assertion). If it fails, surface the new failure — don't try a second production change without the user.

9. **Run targeted regression** — run the changed file's test module + closest siblings. Surface pass/fail counts. Confirm no green-to-red elsewhere.

10. **Route to review** — offer to hand off to `sumo-qa-reviewing-before-merge`. Don't claim "safe to merge" from this skill.

## Special cases

**Retrospective regression — the bug is already fixed.** When the fix landed before the test (hotfix, someone else's fix, a bug spotted after patching), the Iron Law still holds: an unfailed test proves nothing. Manufacture the red against the OLD code:
1. Write the regression test against the *current* (fixed) tree per steps 3–5.
2. Temporarily restore the pre-fix file with a **scoped, reversible** move, **paired with the CORRECT reverse** (each reverse differs; both must end on the current tree). Prefer the first — it leaves the index clean:
   - `git show <pre-fix-commit>:<path> > <path>` → reverse `git checkout -- <path>`.
   - `git checkout <pre-fix-commit> -- <path>` → reverse `git checkout HEAD -- <path>` (or `git restore --source=HEAD --staged --worktree <path>`): it writes the OLD file into BOTH index AND worktree, so a bare `git checkout -- <path>` restores the buggy version from the polluted index — reverse against HEAD, not the index.

   Never destructive — `git reset --hard`, `git checkout <branch>`, `git clean` discard the worktree. Run the test (no command tool: restore nothing, see item 4); capture the real assertion failure.
3. **MANDATORY: reverse the restore (with the reverse paired in step 2) and return to the current worktree before final verification**, then re-run and confirm green. Leaving the tree on the old version reintroduces the bug — the cycle is write-current → restore-old → red → restore-current → green, all BEFORE any handoff; you always end on the current tree. If the pre-fix version is unrecoverable, say so — never fabricate red by weakening the assertion against current code.
4. **Ending:** if not yet run, state each command as a plan (never ask the user to run them, never narrate running them or show their output: step 7's no-run rule applies) and end on step 7's second phrasing; after a real red and green the fix already exists, so skip to step 9.

**Tests for code outside the coverage target.** A test for code outside `--cov=src/sumo_qa` (a script, CI helper, build artifact, sibling package) still counts: coverage doesn't decide whether a regression matters. Write it where the code lives; don't drop it, don't widen `--cov`.

**Closest-sibling selection for a novel test file.** No obvious twin? Copy the closest sibling, most-specific first: (1) **same test pattern**, the same *kind* of test (a loader test for a new loader; pattern beats location); (2) **same target area** (module/subsystem/dir); (3) **same fixture/subprocess style**. Read the match first and mirror its imports, fixtures and assertions.

## Process Flow

The Checklist above.

## Red Flags — STOP and rework

| Thought | Reality |
|---|---|
| "I'll write the test AND the production code in one turn" | Iron Law violated. Red phase first, separately, with evidence shown. |
| "I'll write the test idea AND the test in the same message" | HARD-GATE. Test idea → confirm → write → run; the gate catches misaligned assertions. |
| "Test passed on first run — must already be correct" | The test is wrong; it's not testing what you think. Tighten the assertion until you can see it fail. |
| "Failed with import / syntax / fixture error — that's red" | Not red. A red test fails on its assertion, not on a precondition. |
| "Assertion: `assert add(2,3) == 2+3`" | Tautology. The broken code passes this too. Pick an outcome the bug changes. |
| "February has 28 days, so `anchorDay=31` lands on the 28th" | Recall, not derivation. Feb 2024 has 29 days; Feb 2023 has 28. Trace input → rule → expected for THIS input. Same trap: *"UTC offset is +0"*, *"ASCII is 7-bit"* — year/locale/encoding-dependent. A broken impl that hardcodes the same generic passes the assertion. |
| "I'll stub the prod function with `return total * 0.9` so the test fails meaningfully" | Iron Law violated via the stub. Red-phase stubs are signature-only (`pass` / `raise NotImplementedError`); the 0.9 belongs in the user's green phase. |
| "Mutation testing fits here" | Wrong skill. Mutation follow-up is `sumo-qa-strengthening-tests`. |
| "The bug's already fixed, so I'll commit the test green" | An unred test proves nothing. Manufacture red against the pre-fix version with a scoped reversible restore (`git show <commit>:<path> > <path>` → reverse `git checkout -- <path>`; never `reset --hard` / `checkout <branch>`, and not `git stash` — with the fix committed it reverts to the fixed tree, not the pre-fix file), see red, then reverse (pair the reverse to the restore — `git checkout <commit> -- <path>` pollutes the index, so reverse it against HEAD) and see green. Always end on the current tree. |
| "No shell or a denied run: hand it to the user or a subagent, trace red, write the fix" | A traced failure isn't red, even inline (`AssertionError ... at line N`): test only, then step 7's second phrasing. A denial you didn't hit isn't one: attempt the run; a failed command: fix it, re-run. |

## Examples

### Good (explore, confirm idea, write)

> **User:** *"Fix the VIP-double-discount bug regression-first."*
>
> **AI (homework):** finds `pricing/discount_calculator.py:apply_discounts()` + test file `pricing/test_discount_calculator.py` (pytest, `Order` fixtures); applies decision tables for the VIP/promo combination; proposes `apply_discounts(order_with_vip_and_promo)` asserting `order.total == 90.0`.
>
> **AI (confirm):** *"Test will assert `order.total == 90.0`. **Is 90.0 right — does VIP stack with promo or override it?**"* → **User:** *"VIP overrides promo."* → **AI:** writes the test, runs it, captures `AssertionError: assert 80.0 == 90.0` at `test_discount_calculator.py:47` verbatim, hands off for the green change.

### Good (stateful mock)

> **User:** *"Fix the async memoizer rejection-caching bug regression-first. If the wrapped function rejects, that failure must not poison the cache; a later call for the same key should retry."*
>
> **AI (proposed test idea):** finds `src/cache/memoizeAsync.js:memoizeAsync()` + test file `test/cache/memoizeAsync.test.js` (jest). *"Function `memoizeAsync`; input is a `jest.fn` that rejects ONCE then resolves on retry; assert `memoized(key)` rejects on the first call and resolves to the success value on the second (with `fn` called twice). Technique: state transition testing (no-entry → rejected → retried-then-resolved). Setup-discriminator: the mock must reject-once-then-resolve, NOT always-reject — broken impl caches the rejection (second await rejects); fixed impl retries (second await resolves)."*
>
> **AI (writes test):**
> ```js
> test('does not cache rejected promises; retries on next call', async () => {
>   const fn = jest.fn()
>     .mockRejectedValueOnce(new Error('Failed'))
>     .mockResolvedValue('value:a');
>   const memoized = memoizeAsync(fn);
>
>   await expect(memoized('a')).rejects.toThrow('Failed');
>   await expect(memoized('a')).resolves.toBe('value:a');
>   expect(fn).toHaveBeenCalledTimes(2);
> });
> ```
>
> Discriminator: broken `memoizeAsync` caches the rejection, so the second `await` re-rejects (red); fixed re-invokes `fn` and resolves to `'value:a'` (green). The reject-once-then-resolve mock is what makes broken and fixed diverge.

### Bad (test + prod fix together)

> **AI** (same VIP bug): edits `apply_discounts()` to short-circuit on `customer_tier=VIP`, adds an assertion, runs the suite, reports green.
>
> Iron Law violated. No red phase, no proof the test catches the bug — the assertion may pass on broken code too.

### Bad (recalled expected value, not derived from the specific input)

> **User:** *"Regression test for `nextChargeDate(date, anchorDay)` — Jan 31 anchored to 31 should land on the last valid day of Feb."*
>
> **AI:** *"Assert `nextChargeDate(LocalDate.of(2024, 1, 31), 31) == LocalDate.of(2024, 2, 28)`. Technique: boundary value analysis."*
>
> Expected-value derivation fails: 2024 is a leap year, so Feb 2024's last day is the 29th. The candidate recalled *"Feb has 28 days"* instead of computing from this input — a broken impl hardcoded to Feb=28 passes the assertion. Fix: trace first (*"Jan 31, anchor 31 → Feb 2024 has 29 days → expected 2024-02-29"*), then pin `LocalDate.of(2024, 2, 29)`. Stronger: parametrise across `2024-01-31 → 2024-02-29` AND `2023-01-31 → 2023-02-28` so leap-year clamping becomes its own discriminator.

## Next skill in the chain

Green confirmed + targeted regression passes → `sumo-qa-reviewing-before-merge` for the safe-to-merge verdict against fresh evidence. If part of a rollout dispatched by `sumo-qa-executing-qa-rollout` → `sumo-qa-finishing-qa-work` instead, to capture evidence and produce the PR-ready summary.

Attribution

sumithrsumithr
View sourceSee grades on GitHubMore from sumithr →
SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments (0)

No comments yet. Be the first to comment!

SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Related Skills

Screen Reader Testing

Practical guide to testing web applications with screen readers for comprehensive accessibility validation.

401991 votes

Tdd Workflow

在编写新功能、修复错误或重构代码时使用此技能。强制执行测试驱动开发,包含单元测试、集成测试和端到端测试,覆盖率超过80%。

2456590 votes

Eval Harness

克劳德代码会话的正式评估框架,实施评估驱动开发(EDD)原则

2456590 votes

Python Testing

使用pytest、TDD方法、夹具、模拟、参数化和覆盖率要求的Python测试策略。

2456590 votes

Django Tdd

Django测试策略,包括pytest-django、TDD方法论、factory_boy、模拟、覆盖率以及测试Django REST Framework API。

2456590 votes
View all in testing →