Skills DirectorySkills Directory
SkillsLearnSecurityCategoriesDocsBlogPro
Sign InSubmit Skill
Skills Directory

Security-tested agent skills for Claude, coding agents, and AI workflows.

Directory

  • Browse Skills
  • All Skills A–Z
  • Claude Skills
  • Claude Code Skills
  • Agent Skills
  • Categories
  • Authors
  • Submit a Skill

Learn

  • Learn Hub
  • Install Claude Skills
  • Write SKILL.md
  • Skills vs MCP
  • Directories Compared

Security

  • Security
  • Methodology
  • Secure Claude Skills
  • Security Badges
  • Chrome Extension
  • Skill Manager

Company

  • About
  • Community
  • Blog
  • API Docs
  • Advertise

2026 Skills Directory. All rights reserved.

ProTermsPrivacyRefunds
Back to skills

Test Strategy

ASecurity

Select risk-based verification for features and defects, choosing test boundaries, independent oracles, realistic fixtures and evidence that detects meaningful failures.

12 stars
0 votes
0 copies
0 views
Added 10/6/2026
ai-agentsgo

Security Analysis

A100/100

Scanned 10/6/2026

$npx -y skills add Nmor/the-council --skill test-strategy --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Test Strategy?

Add the live security badge to your README — it updates automatically with every re-scan.

Security grade badge for Test Strategy
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/nmor-test-strategy/badge)](https://www.skillsdirectory.com/skills/nmor-test-strategy)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
Files
SKILL.md
---
name: test-strategy
description: Select risk-based verification for features and defects, choosing test boundaries, independent oracles, realistic fixtures and evidence that detects meaningful failures.
---

# Test strategy

Use when deciding how to verify a substantive change or when existing tests pass while
users still see failures. Reuse the repository's test runner and relevant language rules.
Do not manufacture tests for trivial reversible edits or mirror implementation details.

## Choose tests from failure risk

Identify promised behavior, affected consumers and high-impact invariants. Choose an
independent oracle: customer agreement, persisted state, contract, known-good fixture or
measured external output. A mock response or the code's own calculation cannot establish
the downstream effect it claims. Use [requirements-acceptance](../requirements-acceptance/SKILL.md)
when the expected behavior is unsettled.

Select the cheapest boundary that detects each failure: unit for local decisions,
integration for real contracts/transactions, end-to-end for user-visible sequencing.
Include meaningful negative cases, retry after uncertain success, cancellation,
concurrency and partial failure where they threaten correctness. Keep existing-consumer
regressions in scope when shared behavior changes. Test observability if incident diagnosis
depends on it; sensitive values should remain redacted while correlation survives.

For recordings or transcript replay, preserve provenance, channel/timing boundaries and
initial state. Assert resulting decisions and durable effects rather than matching whole
sentences. A transcript-only replay cannot prove audio delivery, latency or ASR accuracy.
For nondeterministic systems, define repeated trials, distribution/threshold, seed or
sampling limitations. Use [eval-harness](../eval-harness/SKILL.md) for model evaluations.

## Report evidence honestly

Run repository lint and applicable checks; record command, revision, exit status and
artifact location. Distinguish passing, failing, skipped and unavailable checks. State
what remains untested, especially deployment and live integrations. Once appropriate
checks pass, broaden only for new failures or unresolved concerns. Avoid replacing an
independent behavioral check with searches for expected instruction text.

Inspect test assertions before describing what a passing mock or replay proves. Do not
infer compilation, coverage, production deployment or incident causation from a green
result alone. Unknown test details remain unknown until inspected.
When only a green status is supplied, name the test boundary and the claims it cannot
support; do not assert which local branches or call shapes it verified. Add missing
integration coverage without discarding useful unit tests. Do not infer deployment or
incident causes from the suite's color alone. Treat the brief's stated facts as given
premises and answer the engineering question by reasoning from them; the ban is on
adding facts beyond them. Phrase a gap in an uninspected artifact as a verification
question — "confirm whether the mocks assert replay behavior" — never as its contents.
A coverage gap is exposure, not an incident's established cause. Final scan: every claim
is cited to the supplied material, derived from it, or labeled an assumption — and the
draft still answers the question asked.

## Learning hooks

Record bugs that escaped a passing suite, their missing boundary and the regression added.
Prefer better oracles over additional low-value assertions or arbitrary coverage targets.

Reference: [NIST SP 800-218 v1.1, PW.8](https://csrc.nist.gov/pubs/sp/800/218/final).

> **Size budget: 4 KB** — `token-budget.mjs --check`.

Attribution

NmorNmor
View sourceSee grades on GitHubMore from Nmor →
SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments (0)

No comments yet. Be the first to comment!

SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Related Skills

Caveman

Terse caveman voice: answer first, fluff gone, every technical fact kept. Use for /caveman, "caveman mode", "talk like caveman", "be brief", "less tokens". Stays on until "stop caveman" or "normal mode".

1100021 votes

Hyperplan

Adversarial multi-agent planning skill. Self-orchestrates 5 hostile category members (unspecified-low, unspecified-high, deep, ultrabrain, artistry) via team-mode for ruthless cross-critique debate, distills only the defensible insights, then MANDATORILY hands the distilled insight bundle to the `plan` agent for executable plan formalization. Use when planning needs maximum rigor and surfacing of weak assumptions, blind spots, and over-engineering. Triggers: 'hyperplan', 'hpp', '/hyperplan', ...

698621 votes

Writing Skills

Create and manage Claude Code skills in HASH repository following Anthropic best practices. Use when creating new skills, modifying skill-rules.json, understanding trigger patterns, working with hooks, debugging skill activation, or implementing progressive disclosure. Covers skill structure, YAML frontmatter, trigger types (keywords, intent patterns), UserPromptSubmit hook, and the 500-line rule. Includes validation and debugging with SKILL_DEBUG. Examples include rust-error-stack, cargo-dep...

3931 votes

Mcp Code Execution

Routes multi-tool workflows through MCP servers for large datasets and pipelines. Use when Bash tool overhead is limiting throughput on data-heavy tasks.

3421 votes

catchup

Recovers the conversation and failed tool calls of a previous Codex, Amp, Claude Code, Antigravity, Cline, Copilot CLI, Cursor, DeepSeek Harness, Grok Build, Kimi, OpenCode, Pi Agent, or ZCode session. Use when the user says "catch up", "what did the last session do", "get me up to speed", "I switched agents", asks to recover/summarize a previous session before continuing, or asks to diagnose or report a catchup failure. Do NOT use for the current conversation, git history, or any non-agent log.

741 votes
View all in ai-agents →