Skills DirectorySkills Directory
SkillsLearnSecurityCategoriesDocsBlogPro
Sign InSubmit Skill
Skills Directory

Security-tested agent skills for Claude, coding agents, and AI workflows.

Directory

  • Browse Skills
  • All Skills A–Z
  • Claude Skills
  • Claude Code Skills
  • Agent Skills
  • Categories
  • Authors
  • Submit a Skill

Learn

  • Learn Hub
  • Install Claude Skills
  • Write SKILL.md
  • Skills vs MCP
  • Directories Compared

Security

  • Security
  • Methodology
  • Secure Claude Skills
  • Security Badges
  • Chrome Extension
  • Skill Manager

Company

  • About
  • Community
  • Blog
  • API Docs
  • Advertise

2026 Skills Directory. All rights reserved.

ProTermsPrivacyRefunds
Back to skills

Skill Creation

ASecurity

Use when creating a new skill, improving an existing skill, or deciding what a skill should contain and how it should be structured

4 stars
0 votes
0 copies
0 views
Added 9/23/2026
ai-agentsgogitsecurity

Security Analysis

A100/100

Pro scans all 3 files and shows the line behind each finding

Scanned 9/23/2026

$npx -y skills add metraton/gaia --skill skill-creation --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Skill Creation?

Add the live security badge to your README — it updates automatically with every re-scan.

Security grade badge for Skill Creation
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/metraton-skill-creation/badge)](https://www.skillsdirectory.com/skills/metraton-skill-creation)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
Files
SKILL.md
---
name: skill-creation
description: Use when creating a new skill, improving an existing skill, or deciding what a skill should contain and how it should be structured
---

# Skill Creation

## What is a skill?

Injected procedural knowledge -- the "how" for agents. The agent brings identity and domain knowledge. The skill brings process and protocol. They never duplicate each other.

## Step 1: Choose the type

Type determines structure. Choose before writing anything.

| Type | Purpose | When it applies |
|------|---------|-----------------|
| **Discipline** | Enforces rules the agent will rationalize around under pressure | command-execution, execution |
| **Technique** | How to think about or approach a class of problem | investigation, approval |
| **Reference** | Lookup tables, classifications, format specifications | security-tiers, fast-queries, git-conventions |
| **Domain** | Project-specific patterns for a technical area | gaia-patterns |
| **Protocol** | System operating contract -- state machines, mandatory formats | agent-protocol |

## Step 1.5: Situate the skill in its flow

With the type chosen, place the skill before you structure it: where it lives in a flow, and what it can affect. Treat these as one gate with two coupled facets -- position is what sets blast radius.

- **Position.** Is the skill *standalone* (a self-contained process run from a clean start), or does it run *mid-flow* -- sometimes after many skills have already executed -- with an upstream that hands it state and a downstream that consumes what it emits?
- **Blast radius.** What does the skill's output touch or trigger? A standalone technique's reach ends at its own result; a mid-flow skill's is amplified, because its output becomes the next step's input.

The consequence is why the gate exists. A mid-flow skill inherits state and assumptions from upstream; write it as though it started clean and it either redoes work already settled or emits what the next step cannot consume -- and because that output is consumed downstream, the break propagates instead of staying local. A wrong position mis-scopes everything the skill's process assumes and produces.

The gate is **conditional on the type from Step 1**, not universal:

- **Protocol** -- load-bearing: position is the substrate, since a protocol cannot sequence its state machine without knowing where in the larger flow it stands, what the prior turn settled, and what the next needs.
- **Technique / Domain** -- relevant only when the skill runs inside a pipeline; skip it for a self-contained one.
- **Reference** -- irrelevant: a lookup table has no position in a flow. Forcing this gate on a pure Reference skill is itself the "generic without consequence" anti-pattern.

When position is load-bearing **and** the flow is not determinable from the context you were given, ask the user where the skill sits before writing -- do not guess; when the context fixes it or the type makes it irrelevant, do not ask.

This is *frame before action* applied to the skill you are placing: a unit of work has a past that set it up and a future that consumes it. See `agent-protocol` ("Frame before action") for that principle at the agent-turn level.

## Step 2: Apply the type structure

**Discipline:** Iron Law -> Mental Model -> Rules -> Traps -> Anti-patterns. Each trap and anti-pattern names a *principle of failure* -- one row per failure mode, not one row per concrete instance. If three rows are subcases of the same principle, one row that names the principle teaches more than three rows that enumerate.

**Technique:** Overview (core principle + when to use) -> Process (numbered steps) -> Anti-patterns.

**Reference:** Quick-scan table at top -> Examples -> Edge cases / special rules.

**Domain:** Conventions (naming, structure) -> Examples/snippets -> Key rules -> links to reference files.

**Protocol:** State machine / flow -> Mandatory format -> State transitions -> Error handling.

## Step 2.5: Open self-contained

The first sentence states what the skill IS and does, in its own terms. Never open by contrast ("this is NOT X"): that forces the reader to already know X, so the skill stops being self-contained. If you must disambiguate from a sibling skill, do it *after* the self-definition and frame it as a pointer for continuation ("for the universal envelope see agent-protocol"), not as the definition.

Cross-referencing another skill for *flow continuation* (handoff, "see X for the next step") is correct and avoids duplication. Defining your skill *by contrast* with another ("unlike X, this...") is not -- the first keeps you self-contained, the second couples your meaning to a skill the reader may not have loaded.

See `examples.md` for a before/after of each prose failure mode.

## Step 3: Write for judgment, not compliance

A rule without context ("ALWAYS do X") carries almost no weight in the LLM's reasoning -- the model has no reason to prioritize it over competing signals. An explanation with consequences carries enough weight to influence decisions even under pressure. Every line competes for attention; earn each one with reasoning the model can use.

The dual test: for each row in a Traps or Anti-patterns table, ask -- is this a separate principle, or is it the same principle as a row already there? If the latter, fold it in. Specificity competes with itself: five rows that share a principle each carry one-fifth the weight of one row that names the principle, because the model spreads attention across them.

The test: for each rule, ask -- if the agent saw enough examples of this going wrong, would it reach the same conclusion? If yes, you are capturing genuine wisdom. If no, it needs more context.

For detailed guidance on tone by type, see `reference.md`.

## Step 4: Write the description field

Triggering conditions only -- describing the process causes the agent to follow the description and skip reading the content. See `agent-creation/SKILL.md` Step 4 for the bad/good example pattern.

## Step 5: Respect the line budget

| Injection method | Budget | Reason |
|-----------------|--------|--------|
| Frontmatter (always loaded) | < 100 lines | Loaded on every agent call |
| On-demand (read from disk) | < 500 lines | Loaded only when explicitly needed |

Size is not the only criterion. The canonical contract/schema that IS the skill's purpose belongs in SKILL.md even when large; `reference.md` holds the *deep mechanics* (internals, walkthroughs, edge cases) that a reader needs only occasionally. Ask "is this the primary thing the skill exists to state?" before "is this big?".

The deciding criterion is the READER. What someone holding the idea needs in order to reason with it belongs in SKILL.md; what only an implementer needs -- field-by-field schemas, build cycles, invariant tables -- belongs in `reference.md`. Serving both readers from one file is what makes a skill fail to teach: one that mixed them produced authors using 10% of its vocabulary, and splitting them cut it from 404 to 274 lines while an eval showed it taught more.

Heavy reference material -> `reference.md` (on-demand). Concrete examples -> `examples.md`. Executable tools -> `scripts/`.

```
skill-name/
├── SKILL.md          <- main content (always loaded)
├── reference.md      <- heavy docs (on-demand)
├── examples.md       <- concrete examples (on-demand)
└── scripts/          <- executable tools
```

## Step 6: Verify it teaches

Every step above tests the writing; none tests whether a reader learns. Fix the rubric BEFORE seeing any answer -- written afterwards it only rationalizes what you got. Then hand a fresh agent a vague prompt in the real reader's voice, naming neither the skill nor any tool: if it finds the skill unprompted the trigger works, and if it reaches the result without being handed a tool, the skill taught. Where a previous version exists, measure against its readers -- the question is not "did it answer well?" but "does it beat the baseline?".

The rubric must be able to fail in both directions, and both directions pay: one eval exposed a section no reader ever opened (fixed by branching the first read), and another falsified the author's own hypothesis -- a "missing mode" a fresh agent derived unaided, which would otherwise have shipped as an invented section.

For a Domain skill describing a real system, add the coverage test: extract what the system actually does from the code, then check that each mechanic is a consequence of some stated principle. A mechanic no principle explains means a principle is missing, not a row -- that test took one skill from 7 principles to 9. See `reference.md` for how to build the prompt, the baseline, and the coverage extraction.

## When to create vs update

**Create new skill:** Distinct behavioral concern not covered by existing skills. Domain knowledge inline in an agent that applies to multiple agents.

**Update existing skill:** Agent ignores a rule the skill already defines -> strengthen with traps. Skill is missing a type-appropriate section.

**Put elsewhere:** Project-specific config -> CLAUDE.md or agent inline. Single-agent-only behavior -> keep inline. Knowledge the LLM covers well from training -> not needed.

**When creating a new skill:** Also update `skills/README.md` to add the new skill to the index. Load Skill('readme-writing') to do this correctly.

## Anti-Patterns

- **Description summarizes process** -- agent follows the description and skips reading the skill body.
- **Discipline without traps OR with too many traps** -- agents rationalize around bare rules; agents also dilute their attention across overlapping traps. A trap row earns its place by capturing a failure mode the existing rows do not. If you can describe a new trap by adjusting the wording of an existing trap, the existing row is what needed strengthening, not a new neighbor.
- **Generic without consequence** -- "be careful with commands" teaches nothing because it has no consequence attached. The fix is naming what goes wrong when you violate the rule, not enumerating instances of the rule. A principle with one consequence outweighs five examples without one.
- **Duplicates agent content** -- two sources of truth both become stale; pick one place.
- **Single responsibility violated** -- if a skill covers two distinct behaviors, split it.
- **Opening by negation** -- defining the skill by what it is not ("this is NOT agent-protocol"). The reader must already hold the other concept to parse yours. State what it IS first; disambiguate after.
- **Phantom / unanchored reference** -- naming a field, tag, or module that does not exist by that name in code, OR anchoring to line numbers that drift on every edit. Anchor to symbols, not line ranges (`approval_grants.py:1679-1955`) -- a symbol survives edits, a line number does not. For Reference skills especially, every named artifact must be verifiable -- cite the file and symbol. Verify before you assert. And write the anchor in the form the drift checker reads: `path/to/file.py::symbol`, which `tests/layer1_prompt_regression/test_skill_reference_integrity.py::check_text` resolves. A symbol mentioned in prose parentheses is invisible to it, so the citation style alone decides whether a false claim is caught -- see the next row.
- **Quoted implementation** -- pasting another file's body into a skill. The paste is not the implementation, it is a *claim* about it, and a claim with no invalidation trigger: the commit that falsifies it edits the other file and has no reason to visit the skill, so it rots in place while still reading as authoritative. Measured -- the consent adapter quoted a `break` loop that was removed the same day the skill was written, and the skill kept asserting the old semantics against a sibling that had been updated. Name the symbol and state the contract in prose; the symbol's docstring owns the detail, and the reader who needs the body opens it there. An interface example the skill itself OWNS is legitimate -- a call shape, a label format, an envelope the skill defines -- because the skill is that fact's source and nothing else can falsify it. A copy of another file's implementation never is. `skills/code-standards` already forbids this shape for code comments ("Comment that points outward"); this extends the same norm to the substrate agents actually load.
- **Symbol claimed in prose instead of anchored** -- writing a symbol reference as parenthetical prose, `_is_protected`, rather than as the explicit `file::symbol` anchor the drift check matches. The two read identically to a human and differently to the machine: the anchored form is resolved against the file and fails loudly when the symbol is gone, the prose form is not scanned at all. So an unanchored claim is not merely weaker evidence -- it is a claim with no invalidation trigger whatsoever, the same defect as a quoted implementation wearing a citation's clothes. Measured: one deleted symbol left three stale citations across three skills and the checker caught exactly ONE, purely because that one happened to be written as an anchor. Any claim that a symbol EXISTS or that it BEHAVES a certain way carries the anchored form. Prose parentheses are for a symbol the skill is naming in passing, not resting an assertion on.
- **Single-case instruction presented as universal** -- an instruction written from one case reads as covering every case, so the reader of the other case obeys a wrong instruction correctly. One skill paid for this three times: "read `data/` first" (right when modifying, false when creating, where no `data/` exists yet), a cycle that started at "vague idea" with no door for the most frequent entry, and a crucial warning that lived in only one of its documents. If the process has more than one entry, the first instruction branches by case -- name the doors.
- **Inflated prose** -- restates the heading or hedges without adding a decision the reader can act on. If removing the sentence does not change what the agent does, cut it.

Attribution

metratonmetraton
View sourceSee grades on GitHubMore from metraton →
SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments (0)

No comments yet. Be the first to comment!

SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Related Skills

Caveman

Terse caveman voice: answer first, fluff gone, every technical fact kept. Use for /caveman, "caveman mode", "talk like caveman", "be brief", "less tokens". Stays on until "stop caveman" or "normal mode".

1100021 votes

Hyperplan

Adversarial multi-agent planning skill. Self-orchestrates 5 hostile category members (unspecified-low, unspecified-high, deep, ultrabrain, artistry) via team-mode for ruthless cross-critique debate, distills only the defensible insights, then MANDATORILY hands the distilled insight bundle to the `plan` agent for executable plan formalization. Use when planning needs maximum rigor and surfacing of weak assumptions, blind spots, and over-engineering. Triggers: 'hyperplan', 'hpp', '/hyperplan', ...

698461 votes

Writing Skills

Create and manage Claude Code skills in HASH repository following Anthropic best practices. Use when creating new skills, modifying skill-rules.json, understanding trigger patterns, working with hooks, debugging skill activation, or implementing progressive disclosure. Covers skill structure, YAML frontmatter, trigger types (keywords, intent patterns), UserPromptSubmit hook, and the 500-line rule. Includes validation and debugging with SKILL_DEBUG. Examples include rust-error-stack, cargo-dep...

3931 votes

Mcp Code Execution

Routes multi-tool workflows through MCP servers for large datasets and pipelines. Use when Bash tool overhead is limiting throughput on data-heavy tasks.

3421 votes

catchup

Recovers the conversation and failed tool calls of a previous Codex, Amp, Claude Code, Antigravity, Cline, Copilot CLI, Cursor, DeepSeek Harness, Grok Build, Kimi, OpenCode, Pi Agent, or ZCode session. Use when the user says "catch up", "what did the last session do", "get me up to speed", "I switched agents", asks to recover/summarize a previous session before continuing, or asks to diagnose or report a catchup failure. Do NOT use for the current conversation, git history, or any non-agent log.

741 votes
View all in ai-agents →