Skills DirectorySkills Directory
SkillsLearnSecurityCategoriesDocsBlogPro
Sign InSubmit Skill
Skills Directory

Security-tested agent skills for Claude, coding agents, and AI workflows.

Directory

  • Browse Skills
  • All Skills A–Z
  • Claude Skills
  • Claude Code Skills
  • Agent Skills
  • Categories
  • Authors
  • Submit a Skill

Learn

  • Learn Hub
  • Install Claude Skills
  • Write SKILL.md
  • Skills vs MCP
  • Directories Compared

Security

  • Security
  • Methodology
  • Secure Claude Skills
  • Security Badges
  • Chrome Extension
  • Skill Manager

Company

  • About
  • Community
  • Blog
  • API Docs
  • Advertise

2026 Skills Directory. All rights reserved.

ProTermsPrivacyRefunds
Back to skills

Test

ASecurity

Use when the user wants to know if something works — the answer requires running code, not analyzing it. Output is a verdict backed by evidence: passed, failed, or broken. Primary triggers: 'run the tests', 'does X still work after my change?', 'did the merge break anything?', 'verify the fix worked', 'check if the endpoint returns X', 'confirm nothing regressed', 'run tests/unit/test_foo.py', 'let me know the results', 'make sure my changes didn't break anything'. Hard stops — do NOT use for...

119 stars
0 votes
0 copies
2 views
Added 5/28/2026
ai-agentsgorailsdebugginggitapidatabasefrontendbackend

Works with

cliapi

Security Analysis

A100/100

Pro scans all 5 files and shows the line behind each finding

Scanned 5/28/2026

$npx -y skills add avibebuilder/claude-prime --skill test --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Test?

Add the live security badge to your README — it updates automatically with every re-scan.

Security grade badge for Test
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/avibebuilder-test/badge)](https://www.skillsdirectory.com/skills/avibebuilder-test)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
Files
SKILL.md
---
name: test
description: "Use when the user wants to know if something works — the answer requires running code, not analyzing it. Output is a verdict backed by evidence: passed, failed, or broken. Primary triggers: 'run the tests', 'does X still work after my change?', 'did the merge break anything?', 'verify the fix worked', 'check if the endpoint returns X', 'confirm nothing regressed', 'run tests/unit/test_foo.py', 'let me know the results', 'make sure my changes didn't break anything'. Hard stops — do NOT use for: reviewing test code for quality/coverage gaps, debugging why test infrastructure/databases/seed scripts are misbehaving, writing or fixing tests, diagnosing root causes of unexpected behavior. The deciding question: is the user asking for the result of executing something, or asking for help understanding/analyzing/improving something? If it's the latter, use diagnose or review-code instead."
argument-hint: what-to-test-and-outcome
---

## Route fast

| Situation | Steps |
| --- | --- |
| `run tests` or a specific test path | 3 → 5 → 6 |
| The verification claim is already clear from the request or recent context | 4 → 5 → 6 |
| Behavioral claim but no tests exist | 1 → 4 manual path → 5 → 6 |
| Vague claim or broad change | Full flow |

Skip steps whose answers are already known from the conversation.

## Flow

### 1. Define the claim

Figure out what must be true before running anything.

Possible sources:
- explicit argument
- stated request or acceptance criteria
- recent conversation
- recent diff or git status as fallback

Good claim: `expired tokens return 401`.
Bad claim: `the app works`.

If the claim is mushy, sharpen it before you touch the tools. Weak claims create noisy verification.

### 2. Scope only when needed

Skip if the claim is already scoped.

Use the change set to identify:
- **deliverables** — what must hold true
- **coverage** — what existing tests or manual checks could prove each deliverable
- **blast radius** — what nearby behavior could regress

For shared code, config, auth, schema, or similar cross-cutting changes, read `references/regression-strategy.md` before deciding how wide to test.

### 3. Find the project's real test entrypoints

Check in this order:
1. `CLAUDE.md`
2. `package.json`, `Makefile`, `justfile`, `Taskfile`
3. CI config
4. test framework config
5. existing test layout

Prefer project-defined commands over raw framework commands.

### 4. Choose the proof lane

Match the claim to the execution lane.

| Claim type | Lane |
| --- | --- |
| logic, transforms, business rules | direct tests |
| API behavior or contracts | backend verification |
| UI behavior or rendering | frontend verification |
| visual UI criteria | frontend visual verification |
| DB persistence | backend verification |
| CLI behavior | direct command verification |
| type or schema correctness | direct static verification |
| compile/build correctness | direct static/build verification |

Once you've picked a lane, read the relevant reference before executing. Let the reference own the execution details.

## Reference map

- `references/backend-verification.md` — API, DB, services, and background jobs
- `references/frontend-verification.md` — browser, component, and visual verification
- `references/regression-strategy.md` — how wide to test once the main claim is proven

The right method is the one that can disprove the claim fastest without pretending to offer more confidence than it really does.

If no relevant tests exist, use a manual path instead of calling the result inconclusive. Follow the selected lane's reference rather than rebuilding the checklist inline.

For command-line claims, run the command and inspect output directly.

Keep these guardrails in mind:
- once a lane is chosen, follow its reference instead of rebuilding the same checklist inline
- keep verification outcome-first: prove the user-visible or system-visible result, not just an intermediate action
- when regression is warranted, widen in rings: direct dependents → feature area → full suite

Only use **INCONCLUSIVE** when neither an automated path nor a manual path is feasible.

### 5. Pre-flight and execute

Before running verification, confirm required infra is already available. If a server, DB, queue worker, mock service, or migration is needed and not running, stop and tell the user the exact command to start it. Do **not** start it yourself.

When you run verification, capture evidence:
- exact command
- exit code
- relevant output
- non-zero test count when a suite was run
- warnings, skipped tests, and flaky behavior that matter to the claim

Start with the most direct proof. Expand into broader regression only when the blast radius justifies it.

### 6. Report

Use this structure:

```md
## Verification Report

**Claim**: {what was tested}
**Verdict**: CONFIRMED | REFUTED | PARTIAL | INCONCLUSIVE

### Evidence
- {command}: exit {code} — proves or refutes {deliverable}

### Failures
- {test or command}: {error snippet} — {what it means}

### Gaps
- {deliverable or regression area still unverified}

### Suite Stats
Total: X | Passed: X | Failed: X | Skipped: X | Duration: Xs
Coverage: Lines X% | Branches X% (if available)
```

Use these verdicts consistently:
- **CONFIRMED** — every deliverable has passing evidence
- **REFUTED** — at least one deliverable failed with clear evidence
- **PARTIAL** — some deliverables are proven, others remain unverified
- **INCONCLUSIVE** — verification was blocked by missing infra, environment issues, or unavailable access

## Common verification situations

### Bug-fix verification
Prove the original failure no longer reproduces and check one regression ring around the changed area.

### Feature verification
Prove the stated acceptance criteria and include type-check evidence when that meaningfully covers changed contracts.

### Standalone verification
Take the claim from the argument or, if needed, infer it from recent changes.

## Rules

- Run and report only. Do not fix failures.
- A passing suite is not proof if it never exercised the claim.
- Coverage gaps stay in the report even when existing tests pass.
- Flaky behavior and pre-existing failures are findings, not noise.
- If a lane points to a reference, read it before improvising tactics.
- Temporary, verification-only app changes are allowed when they make the proof materially better and the app is ours to edit. Keep them minimal, use them to verify the claim, then remove them before reporting.
- Do not leave verification hooks behind unless the user asked for a durable selector as part of the product change.
- Never start dev servers, databases, builds, or migrations yourself.

## Verification Target

<target>$ARGUMENTS</target>

Attribution

avibebuilderavibebuilder
View sourceSee grades on GitHubMore from avibebuilder →
SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments (0)

No comments yet. Be the first to comment!

SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Related Skills

Caveman

Terse caveman voice: answer first, fluff gone, every technical fact kept. Use for /caveman, "caveman mode", "talk like caveman", "be brief", "less tokens". Stays on until "stop caveman" or "normal mode".

1100021 votes

Hyperplan

Adversarial multi-agent planning skill. Self-orchestrates 5 hostile category members (unspecified-low, unspecified-high, deep, ultrabrain, artistry) via team-mode for ruthless cross-critique debate, distills only the defensible insights, then MANDATORILY hands the distilled insight bundle to the `plan` agent for executable plan formalization. Use when planning needs maximum rigor and surfacing of weak assumptions, blind spots, and over-engineering. Triggers: 'hyperplan', 'hpp', '/hyperplan', ...

698461 votes

Writing Skills

Create and manage Claude Code skills in HASH repository following Anthropic best practices. Use when creating new skills, modifying skill-rules.json, understanding trigger patterns, working with hooks, debugging skill activation, or implementing progressive disclosure. Covers skill structure, YAML frontmatter, trigger types (keywords, intent patterns), UserPromptSubmit hook, and the 500-line rule. Includes validation and debugging with SKILL_DEBUG. Examples include rust-error-stack, cargo-dep...

3931 votes

Mcp Code Execution

Routes multi-tool workflows through MCP servers for large datasets and pipelines. Use when Bash tool overhead is limiting throughput on data-heavy tasks.

3421 votes

catchup

Recovers the conversation and failed tool calls of a previous Codex, Amp, Claude Code, Antigravity, Cline, Copilot CLI, Cursor, DeepSeek Harness, Grok Build, Kimi, OpenCode, Pi Agent, or ZCode session. Use when the user says "catch up", "what did the last session do", "get me up to speed", "I switched agents", asks to recover/summarize a previous session before continuing, or asks to diagnose or report a catchup failure. Do NOT use for the current conversation, git history, or any non-agent log.

741 votes
View all in ai-agents →