Skills DirectorySkills Directory
SkillsLearnSecurityCategoriesDocsCommunityBlog
Sign InSubmit Skill
Skills Directory

Security-tested agent skills for Claude, coding agents, and AI workflows.

Directory

  • Browse Skills
  • All Skills A–Z
  • Claude Skills
  • Claude Code Skills
  • Agent Skills
  • Categories
  • Submit a Skill

Learn

  • Learn Hub
  • Install Claude Skills
  • Write SKILL.md
  • Skills vs MCP
  • Directories Compared

Security

  • Security
  • Methodology
  • Secure Claude Skills
  • Security Badges

Company

  • About
  • Community
  • Blog
  • API Docs
  • Advertise

2026 Skills Directory. All rights reserved.

Back to skills

Behavior Validation

ASecurity

Validates a running application, CLI, API, service, or generated artifact as a user or operator against a prewritten observable behavior contract while remaining source-blind. Use for acceptance checks, runtime proof, anti-fake probes, release smoke tests, or an independent companion to code review. Not for source-quality findings, root-cause diagnosis, or visual design judgment outside the contract.

95 stars
0 votes
0 copies
0 views
Added 9/22/2026
ai-agentsdebugginggitapi

Works with

terminalcliapi

Security Analysis

A100/100

Scanned 9/22/2026

Install to Claude Code

$npx -y skills add thiientv/godmode --skill behavior-validation --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Behavior Validation?

Add the live security badge to your README — it updates automatically with every re-scan.

Security grade badge for Behavior Validation
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/thiientv-behavior-validation/badge)](https://www.skillsdirectory.com/skills/thiientv-behavior-validation)

More formats (shields.io, HTML) on the badges page.

Download Zip
Files
SKILL.md
---
name: behavior-validation
description: >-
  Validates a running application, CLI, API, service, or generated artifact as
  a user or operator against a prewritten observable behavior contract while
  remaining source-blind. Use for acceptance checks, runtime proof, anti-fake
  probes, release smoke tests, or an independent companion to code review. Not
  for source-quality findings, root-cause diagnosis, or visual design judgment
  outside the contract.
---

# Behavior Validation

Judge what the product does, not how its source appears to do it.

## Establish isolation

Write or read the behavior contract before exercising the target. Use
[behavior-contract.md](references/behavior-contract.md). The contract must name
user tasks, expected outcomes, setup, allowed interfaces, negative cases, and
required evidence.

Do not inspect source, diffs, internal tests, Git history, or implementation
notes during validation. Interact only through public browser, CLI, API,
artifact, accessibility, or operator surfaces. If source is required, mark the
clause blocked and hand it to `root-cause-debugging`.

## Exercise the contract

1. Run each task through the same entry point a real user or operator uses.
2. Vary input and state to detect hard-coded success, stale data, or display-only
   behavior.
3. Test invalid, empty, interrupted, retry, persistence, and permission paths
   where the contract makes them relevant.
4. Capture redacted screenshots, terminal excerpts, response summaries, or
   artifact facts.
5. Mark every clause pass, fail, blocked, or out of scope. Lack of evidence is
   not a pass.

When a finding is fixed, rerun the failed clause and nearby regression probes;
do not rerun unrelated expensive scenarios without reason.

## Completion condition

Every relevant contract clause has a status and reproducible evidence, anti-fake
probes were attempted, secrets and private data were excluded, and the report
does not infer implementation defects from observable symptoms.

Attribution

thiientvthiientv
View sourceMore from thiientv →
SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments (0)

No comments yet. Be the first to comment!

SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Related Skills

Caveman

Ultra-compressed communication mode that cuts output tokens while keeping technical accuracy. Levels: lite, full, ultra and the wenyan variants. Use for /caveman, "caveman mode", "talk like caveman", "be brief" or "less tokens".

1066601 votes

Hyperplan

Adversarial multi-agent planning skill. Self-orchestrates 5 hostile category members (unspecified-low, unspecified-high, deep, ultrabrain, artistry) via team-mode for ruthless cross-critique debate, distills only the defensible insights, then MANDATORILY hands the distilled insight bundle to the `plan` agent for executable plan formalization. Use when planning needs maximum rigor and surfacing of weak assumptions, blind spots, and over-engineering. Triggers: 'hyperplan', 'hpp', '/hyperplan', ...

686011 votes

Mcp Code Execution

Routes multi-tool workflows through MCP servers for large datasets and pipelines. Use when Bash tool overhead is limiting throughput on data-heavy tasks.

3351 votes

catchup

Recovers the conversation and failed tool calls of a previous Codex, Claude Code, Antigravity, Cline, Copilot CLI, Cursor, DeepSeek Harness, Kimi, OpenCode, Pi Agent, or ZCode session. Use when the user says "catch up", "what did the last session do", "get me up to speed", "I switched agents", asks to recover/summarize a previous session before continuing, or asks to diagnose or report a catchup failure. Do NOT use for the current conversation, git history, or any non-agent log.

651 votes

math-skill

A comprehensive mathematical reasoning skill for AI assistants — handles arithmetic to research-level problems with rigorous step-by-step reasoning, systematic verification, and transparent uncertainty handling

381 votes
View all in ai-agents →