Skills DirectorySkills Directory
SkillsLearnSecurityCategoriesDocsBlogPro
Sign InSubmit Skill
Skills Directory

Security-tested agent skills for Claude, coding agents, and AI workflows.

Directory

  • Browse Skills
  • All Skills A–Z
  • Claude Skills
  • Claude Code Skills
  • Agent Skills
  • Categories
  • Authors
  • Submit a Skill

Learn

  • Learn Hub
  • Install Claude Skills
  • Write SKILL.md
  • Skills vs MCP
  • Directories Compared

Security

  • Security
  • Methodology
  • Secure Claude Skills
  • Security Badges
  • Chrome Extension
  • Skill Manager

Company

  • About
  • Community
  • Blog
  • API Docs
  • Advertise

2026 Skills Directory. All rights reserved.

ProTermsPrivacyRefunds
Back to skills

Testing Discipline

ASecurity

Use when adding/fixing tests, reproducing bugs, checking limits or failures, or deciding what evidence establishes completion. Covers isolated storage, real domain logic, meaningful boundary regressions and verification of the affected runtime/UI surface. Debugging strategy lives in debug-incident-protocol.

3 stars
0 votes
0 copies
1 views
Added 9/22/2026
ai-agentstestingdebuggingapidocumentation

Works with

api

Security Analysis

A100/100

Pro scans all 2 files and shows the line behind each finding

Scanned 10/6/2026

$npx -y skills add oleg494/coding-kit --skill testing-discipline --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Testing Discipline?

Add the live security badge to your README — it updates automatically with every re-scan.

Security grade badge for Testing Discipline
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/oleg494-testing-discipline/badge)](https://www.skillsdirectory.com/skills/oleg494-testing-discipline)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
Files
SKILL.md
---
name: testing-discipline
description: 'Use when adding/fixing tests, reproducing bugs, checking limits or failures, or deciding what evidence establishes completion. Covers isolated storage, real domain logic, meaningful boundary regressions and verification of the affected runtime/UI surface. Debugging strategy lives in debug-incident-protocol.'
license: MIT
compatibility: pytest, jest and similar; applicable to any language
metadata:
  version: "4.7.0"
---

# Testing discipline: tests as a spec and defining "done"

## 0. TDD gate (OPS §3 Phase 2 companion)

Define observable acceptance and a suitable check before changing behavior. Bug fixes require a failing reproduction followed by a passing result. Preserve meaningful existing tests; keep a new regression when a plausible recurring bug would fail it. Smoke probes and rendered UI interaction are valid evidence, not an obligation to add permanent test files.

## 1. Isolation and structure

1. **NEVER TOUCH PROD STORE IN TESTS** — use throwaway storage selected before import: env override, temporary path and isolated fixture. Do not connect to production merely to compare row counts.
2. **CONTROL IMPORT SIDE EFFECTS** — prefer explicit dependency injection or existing isolated fixtures; use a fresh import when a module actually captures storage at import time.
3. **DOMAIN FIRST** — exercise business rules directly; integration checks cover meaningful I/O boundaries. No arbitrary runtime target or mandatory framework bypass.
4. **FAKE THE EDGES, NOT THE CORE** — replace external Telegram/HTTP/API calls, not the business logic under test. Use a real temporary DB when persistence semantics are part of the behavior.

## 2. Names and boundaries

- **TEST NAMES ARE THE SPEC** — `test_referral_no_self`, `test_crypto_idempotent` — name = rule. `pytest --collect-only` reads like a product checklist. NOT test_1, test_works.
- **PRODUCT RULES AS NAMED TESTS** — the spec lives in tests: `test_free_spent_first`, `test_cannot_buy_while_paid_remains`. A new developer reads the tests = understands the product.
- **ASSERT MATERIAL BOUNDARIES** — choose plausible failure modes: zero/exhausted value, self-referral, double credit, empty input, overflow. No happy/edge/abuse quota for every function.
- **WRITE THE ABUSE CASE WHEN YOU WRITE THE GROWTH CASE** — referral/promo written with an anti-fraud test in the same PR.

## 3. Specific tests

- **MONEY PATH**: double-submit → balance +X not +2X; provider error injection → balance unchanged; reject → `assert not user_exists(...)`.
- **LIMITS (cap)**: test before implementation: cap exhausted → False, user not created; monkeypatch.setenv; `assert cap and invited >= cap`.
- **RATE LIMIT**: two consecutive calls → second blocked without external API: mock API → assert len(calls) <= 1.
- **UI** — exercise the actual rendered user flow, states and responsive behavior affected by the change. Permanent tests defend uncertain behavior, not component presence or routing echoes.
- **COPY** — verify meaning, accessibility and any legal/protocol wording that is genuinely contractual. Ordinary editorial changes do not need substring-lock tests.
- **PAYLOAD VALIDATION** — `assert len(body.encode("utf-8")) <= PLATFORM_LIMIT` before deploy.

## 4. DoD — defining "done"

Completion covers every requested behavior, not merely passing tests.

- Run syntax/import/build checks applicable to the changed implementation.
- Run focused existing tests, broadening for shared code or exposed failures.
- Exercise the changed entrypoint, API or rendered UI when runtime behavior is
  the claim. A startup log alone does not prove the changed path.
- A documentation or read-only task does not require an unrelated live service.
  Missing runtime capability: use an isolated smoke probe where possible and
  state exactly what remains unverified.
- Reuse evidence valid for the final checked state; a chat turn alone does not
  invalidate it. See `verification-before-completion`.

## Workflow

1. Define consumer-visible acceptance and the failure being defended.
2. Select the appropriate layer and isolate external side effects.
3. Reproduce bugs before fixing them; for new behavior define a focused check.
4. Implement the complete request, run the check and applicable existing suite.
5. Verify the affected runtime surface and report the evidence's actual scope.

## Checklist

- [ ] Storage and network effects are isolated from production
- [ ] Checks exercise real domain behavior, boundaries or failure transitions
- [ ] No source-text, wording, mock-echo or test-count padding
- [ ] Money and limits retain idempotency and no-side-effects-on-reject coverage
- [ ] Every acceptance criterion has appropriate observed evidence

Attribution

oleg494oleg494
View sourceSee grades on GitHubMore from oleg494 →
SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments (0)

No comments yet. Be the first to comment!

SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Related Skills

Caveman

Terse caveman voice: answer first, fluff gone, every technical fact kept. Use for /caveman, "caveman mode", "talk like caveman", "be brief", "less tokens". Stays on until "stop caveman" or "normal mode".

1100021 votes

Hyperplan

Adversarial multi-agent planning skill. Self-orchestrates 5 hostile category members (unspecified-low, unspecified-high, deep, ultrabrain, artistry) via team-mode for ruthless cross-critique debate, distills only the defensible insights, then MANDATORILY hands the distilled insight bundle to the `plan` agent for executable plan formalization. Use when planning needs maximum rigor and surfacing of weak assumptions, blind spots, and over-engineering. Triggers: 'hyperplan', 'hpp', '/hyperplan', ...

698461 votes

Writing Skills

Create and manage Claude Code skills in HASH repository following Anthropic best practices. Use when creating new skills, modifying skill-rules.json, understanding trigger patterns, working with hooks, debugging skill activation, or implementing progressive disclosure. Covers skill structure, YAML frontmatter, trigger types (keywords, intent patterns), UserPromptSubmit hook, and the 500-line rule. Includes validation and debugging with SKILL_DEBUG. Examples include rust-error-stack, cargo-dep...

3931 votes

Mcp Code Execution

Routes multi-tool workflows through MCP servers for large datasets and pipelines. Use when Bash tool overhead is limiting throughput on data-heavy tasks.

3421 votes

catchup

Recovers the conversation and failed tool calls of a previous Codex, Amp, Claude Code, Antigravity, Cline, Copilot CLI, Cursor, DeepSeek Harness, Grok Build, Kimi, OpenCode, Pi Agent, or ZCode session. Use when the user says "catch up", "what did the last session do", "get me up to speed", "I switched agents", asks to recover/summarize a previous session before continuing, or asks to diagnose or report a catchup failure. Do NOT use for the current conversation, git history, or any non-agent log.

741 votes
View all in ai-agents →