Skills DirectorySkills Directory
SkillsLearnSecurityCategoriesDocsBlogPro
Sign InSubmit Skill
Skills Directory

Security-tested agent skills for Claude, coding agents, and AI workflows.

Directory

  • Browse Skills
  • All Skills A–Z
  • Claude Skills
  • Claude Code Skills
  • Agent Skills
  • Categories
  • Authors
  • Submit a Skill

Learn

  • Learn Hub
  • Install Claude Skills
  • Write SKILL.md
  • Skills vs MCP
  • Directories Compared

Security

  • Security
  • Methodology
  • Secure Claude Skills
  • Security Badges
  • Chrome Extension
  • Skill Manager

Company

  • About
  • Community
  • Blog
  • API Docs
  • Advertise

2026 Skills Directory. All rights reserved.

ProTermsPrivacyRefunds
Back to skills

Sre

ASecurity

Use when an org role acts as SRE and must make reliability a budgeted feature: user-facing SLOs, error-budget policy, observability and toil automation. Role guidance; for SLO definitions, runbooks and monitoring configs produced as artifacts see sre-engineer.

21 stars
0 votes
0 copies
0 views
Added 9/22/2026
ai-agentsgogitdevops

Security Analysis

A100/100

Scanned 9/28/2026

$npx -y skills add monoes/monomind --skill sre --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Sre?

Add the live security badge to your README — it updates automatically with every re-scan.

Security grade badge for Sre
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/monoes-sre/badge)](https://www.skillsdirectory.com/skills/monoes-sre)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
Files
SKILL.md
---
name: sre
description: "Use when an org role acts as SRE and must make reliability a budgeted feature: user-facing SLOs, error-budget policy, observability and toil automation. Role guidance; for SLO definitions, runbooks and monitoring configs produced as artifacts see sre-engineer."
tags: ["devops","reliability","observability"]
tools: []
license: Apache-2.0
source: https://github.com/monoes/monomind
---
# SRE — Best Practices

## Focus
Treat reliability as a measurable, budgeted feature — define SLOs that reflect real user experience, build observability that answers questions before they're asked, and automate away toil.

## Best practices
- Define SLOs from user-facing behavior (availability, latency) with an explicit target and measurement window — not arbitrary round numbers picked without data.
- Let the error budget drive prioritization: budget remaining → ship features; budget exhausted → reliability work takes priority, no exceptions.
- Build observability across all three pillars — metrics for trends/alerting, logs for event detail, traces for cross-service request flow — so "why is this broken?" has an answer in minutes.
- Automate anything done manually twice; toil that isn't automated compounds as the system scales.
- Roll out changes progressively (canary → percentage → full) and never big-bang deploy to 100% of traffic.
- Set burn-rate alerts (fast burn = page now, slow burn = ticket) rather than a single static threshold.
- Run chaos engineering exercises proactively to find weaknesses before users do, not just after an incident.

## Common pitfalls
- Setting SLO targets that don't map to anything users actually experience (e.g. arbitrary "five nines" with no cost/benefit analysis).
- Doing reliability work without data showing there's a problem — optimizing based on intuition instead of measured burn rate.
- Alerting on every anomaly instead of on SLO burn rate, producing pager fatigue that trains engineers to ignore pages.
- Treating each nine of availability as linearly as expensive as the last — it isn't, and pretending otherwise misallocates effort.
- Fixing incidents by hand repeatedly instead of turning the fix into an automated runbook.

## Tools & techniques
- SLI/SLO/error-budget framework with burn-rate multi-window alerts (e.g. 14.4x/1h for critical, 6x/6h for warning).
- The four golden signals (latency, traffic, errors, saturation) as the baseline dashboard for every service.
- Chaos engineering tooling (fault injection, game days) to validate resilience assumptions under controlled conditions.
- Blameless post-incident review focused on systemic fixes, tracked to completion — not just narrated once and forgotten.

Attribution

monoesmonoes
View sourceSee grades on GitHubMore from monoes →
SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments (0)

No comments yet. Be the first to comment!

SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Related Skills

Caveman

Terse caveman voice: answer first, fluff gone, every technical fact kept. Use for /caveman, "caveman mode", "talk like caveman", "be brief", "less tokens". Stays on until "stop caveman" or "normal mode".

1100021 votes

Hyperplan

Adversarial multi-agent planning skill. Self-orchestrates 5 hostile category members (unspecified-low, unspecified-high, deep, ultrabrain, artistry) via team-mode for ruthless cross-critique debate, distills only the defensible insights, then MANDATORILY hands the distilled insight bundle to the `plan` agent for executable plan formalization. Use when planning needs maximum rigor and surfacing of weak assumptions, blind spots, and over-engineering. Triggers: 'hyperplan', 'hpp', '/hyperplan', ...

698621 votes

Writing Skills

Create and manage Claude Code skills in HASH repository following Anthropic best practices. Use when creating new skills, modifying skill-rules.json, understanding trigger patterns, working with hooks, debugging skill activation, or implementing progressive disclosure. Covers skill structure, YAML frontmatter, trigger types (keywords, intent patterns), UserPromptSubmit hook, and the 500-line rule. Includes validation and debugging with SKILL_DEBUG. Examples include rust-error-stack, cargo-dep...

3931 votes

Mcp Code Execution

Routes multi-tool workflows through MCP servers for large datasets and pipelines. Use when Bash tool overhead is limiting throughput on data-heavy tasks.

3421 votes

catchup

Recovers the conversation and failed tool calls of a previous Codex, Amp, Claude Code, Antigravity, Cline, Copilot CLI, Cursor, DeepSeek Harness, Grok Build, Kimi, OpenCode, Pi Agent, or ZCode session. Use when the user says "catch up", "what did the last session do", "get me up to speed", "I switched agents", asks to recover/summarize a previous session before continuing, or asks to diagnose or report a catchup failure. Do NOT use for the current conversation, git history, or any non-agent log.

741 votes
View all in ai-agents →