Skills DirectorySkills Directory
SkillsLearnSecurityCategoriesDocsBlogPro
Sign InSubmit Skill
Skills Directory

Security-tested agent skills for Claude, coding agents, and AI workflows.

Directory

  • Browse Skills
  • All Skills A–Z
  • Claude Skills
  • Claude Code Skills
  • Agent Skills
  • Categories
  • Authors
  • Submit a Skill

Learn

  • Learn Hub
  • Install Claude Skills
  • Write SKILL.md
  • Skills vs MCP
  • Directories Compared

Security

  • Security
  • Methodology
  • Secure Claude Skills
  • Security Badges
  • Chrome Extension
  • Skill Manager

Company

  • About
  • Community
  • Blog
  • API Docs
  • Advertise

2026 Skills Directory. All rights reserved.

ProTermsPrivacyRefunds
Back to skills

Runbook Writer

ASecurity

Writing runbooks that work at 3am — structure, diagnostics, and safe procedures — use when documenting operational response.

2 stars
0 votes
0 copies
3 views
Added 9/29/2026
ai-agentsgo

Works with

cli

Security Analysis

A100/100

Scanned 9/29/2026

$npx -y skills add aicodedecode/awesome-muse-skills --skill runbook-writer --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Runbook Writer?

Add the live security badge to your README — it updates automatically with every re-scan.

Security grade badge for Runbook Writer
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/aicodedecode-runbook-writer/badge)](https://www.skillsdirectory.com/skills/aicodedecode-runbook-writer)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
Files
SKILL.md
---
name: runbook-writer
description: Writing runbooks that work at 3am — structure, diagnostics, and safe procedures — use when documenting operational response.
category: operations
---

## Overview

A runbook is the difference between a 10-minute mitigation and a 2-hour
adventure: the exact steps to diagnose and fix a known failure mode, written
for a tired human under pressure. This skill covers structuring runbooks,
writing procedures that are safe to follow half-asleep, and keeping them
alive as systems change.

## When to use

- Writing a runbook for a new alert or known failure mode
- Turning tribal knowledge into documented procedures
- Reviewing runbooks for clarity, safety, and freshness
- Linking runbooks to alerts so on-call can act immediately
- Auditing runbook coverage across services

## Core concepts

**Write for the 3am reader.** Short sentences, numbered steps, no assumed
context, exact commands to copy-paste (with placeholders clearly marked like
`<POD_NAME>`). The reader is tired, stressed, and possibly unfamiliar with
this service — clarity beats elegance.

**Structure: symptom → diagnose → mitigate → verify → escalate.** Every
runbook opens with how to recognize the problem (alert name, dashboard link,
symptoms), then diagnostic steps to confirm, then the fix, then how to verify
it's fixed, then when to give up and escalate. This shape is scannable under
pressure.

**Commands must be copy-paste safe.** Provide full commands, not fragments;
mark destructive commands explicitly (`⚠️ destructive: deletes data`);
prefer reversible actions first. Include expected output so the reader can
confirm each step worked — "you should see X" prevents blind continuation
past a failed step.

**One runbook per alert.** The alert fires → the runbook link is right there
→ the responder follows it. Orphan alerts with no runbook, and runbooks with
no alert, both rot. Maintain the 1:1 mapping as alerts change.

**Runbooks are code-adjacent: version and test them.** Keep runbooks in
version control next to the service; review them in PRs that change the
system; and test them — game days and incident simulations are runbook tests.
An untested runbook is a hypothesis.

## Practical workflow

1. **Template every runbook:** title, owner, last-verified date, alert link,
   severity, symptoms, diagnosis steps, mitigation steps, verification,
   escalation contacts, related links.
2. **Write the diagnosis section** as a decision tree: check A → if X do
   step 3, if Y jump to step 5. Branching beats linear guessing.
3. **Write mitigation as numbered copy-paste commands** with expected outputs
   and rollback steps for each destructive action.
4. **Define "done":** exact verification (metric back under threshold for
   10 min, error rate < 0.1%, synthetic check green) — not "looks okay".
5. **Define "escalate":** conditions for stopping (step failed twice,
   symptoms don't match, 30 minutes without progress) and who to call —
   escalation is a procedure, not a failure.
6. **Maintain:** every incident that used (or should have used) a runbook
   updates it; quarterly review of stale ones; delete runbooks for
   decommissioned alerts.

## Common pitfalls

- **Wall-of-text runbooks** — paragraphs of background before the first
  actionable step; put actions first, context in an appendix.
- **Missing expected outputs** — the reader can't tell if a step worked;
  show what success looks like at each step.
- **Untested commands** — flags changed, CLIs updated, paths moved; verify
  commands actually run in the current environment.
- **No escalation path** — runbooks that assume the fix always works leave
  responders stranded; always define the exit.
- **Tribal knowledge never written down** — "ask Priya, she knows" is not a
  runbook; capture it before Priya's vacation.
- **Stale runbooks** — referencing decommissioned dashboards, old hostnames,
  or removed flags; the last-verified date and post-incident updates prevent
  rot.

Attribution

aicodedecodeaicodedecode
View sourceSee grades on GitHubMore from aicodedecode →
SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments (0)

No comments yet. Be the first to comment!

SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Related Skills

Caveman

Terse caveman voice: answer first, fluff gone, every technical fact kept. Use for /caveman, "caveman mode", "talk like caveman", "be brief", "less tokens". Stays on until "stop caveman" or "normal mode".

1100021 votes

Hyperplan

Adversarial multi-agent planning skill. Self-orchestrates 5 hostile category members (unspecified-low, unspecified-high, deep, ultrabrain, artistry) via team-mode for ruthless cross-critique debate, distills only the defensible insights, then MANDATORILY hands the distilled insight bundle to the `plan` agent for executable plan formalization. Use when planning needs maximum rigor and surfacing of weak assumptions, blind spots, and over-engineering. Triggers: 'hyperplan', 'hpp', '/hyperplan', ...

698431 votes

Writing Skills

Create and manage Claude Code skills in HASH repository following Anthropic best practices. Use when creating new skills, modifying skill-rules.json, understanding trigger patterns, working with hooks, debugging skill activation, or implementing progressive disclosure. Covers skill structure, YAML frontmatter, trigger types (keywords, intent patterns), UserPromptSubmit hook, and the 500-line rule. Includes validation and debugging with SKILL_DEBUG. Examples include rust-error-stack, cargo-dep...

3931 votes

Mcp Code Execution

Routes multi-tool workflows through MCP servers for large datasets and pipelines. Use when Bash tool overhead is limiting throughput on data-heavy tasks.

3421 votes

catchup

Recovers the conversation and failed tool calls of a previous Codex, Amp, Claude Code, Antigravity, Cline, Copilot CLI, Cursor, DeepSeek Harness, Grok Build, Kimi, OpenCode, Pi Agent, or ZCode session. Use when the user says "catch up", "what did the last session do", "get me up to speed", "I switched agents", asks to recover/summarize a previous session before continuing, or asks to diagnose or report a catchup failure. Do NOT use for the current conversation, git history, or any non-agent log.

741 votes
View all in ai-agents →