Skills DirectorySkills Directory
SkillsLearnSecurityCategoriesDocsBlogPro
Sign InSubmit Skill
Skills Directory

Security-tested agent skills for Claude, coding agents, and AI workflows.

Directory

  • Browse Skills
  • All Skills A–Z
  • Claude Skills
  • Claude Code Skills
  • Agent Skills
  • Categories
  • Authors
  • Submit a Skill

Learn

  • Learn Hub
  • Install Claude Skills
  • Write SKILL.md
  • Skills vs MCP
  • Directories Compared

Security

  • Security
  • Methodology
  • Secure Claude Skills
  • Security Badges
  • Chrome Extension
  • Skill Manager

Company

  • About
  • Community
  • Blog
  • API Docs
  • Advertise

2026 Skills Directory. All rights reserved.

ProTermsPrivacyRefunds
Back to skills

Senior Devops

ASecurity

DevOps/SRE perspective: CI/CD design, infrastructure as code, observability, incident response, and reliability trade-offs. Use when building pipelines, managing infrastructure, or improving operational maturity.

2 stars
0 votes
0 copies
0 views
Added 9/29/2026
ai-agentsgoterraformtestingdevopsci/cd

Works with

cli

Security Analysis

A100/100

Scanned 9/29/2026

$npx -y skills add aicodedecode/awesome-muse-skills --skill senior-devops --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Senior Devops?

Add the live security badge to your README — it updates automatically with every re-scan.

Security grade badge for Senior Devops
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/aicodedecode-senior-devops/badge)](https://www.skillsdirectory.com/skills/aicodedecode-senior-devops)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
Files
SKILL.md
---
name: senior-devops
description: DevOps/SRE perspective: CI/CD design, infrastructure as code, observability, incident response, and reliability trade-offs. Use when building pipelines, managing infrastructure, or improving operational maturity.
category: development
---

# Senior DevOps Engineer

## Overview

Senior DevOps is **applied reliability**: making shipping fast *and* safe, and making systems
explain themselves when they break. This skill captures that practice — pipeline design, immutable
infrastructure, observability that actually helps at 3am, and incident response without heroics.

The through-line: automate the toil, measure what matters, and design for the failure you haven't
had yet.

## When to use

- Designing or fixing CI/CD pipelines.
- Choosing infrastructure approaches (containers, IaC, managed services).
- Setting up monitoring, alerting, and on-call practices.
- Responding to or learning from production incidents.
- Reducing deploy friction, lead time, or change-failure rate.

## Core concepts

- **Everything as code.** Infrastructure, pipeline definitions, dashboards, and alerts live in
  version control and go through review like application code. Click-ops is unreviewable,
  unrepeatable, and unauditable.
- **Immutable artifacts.** Build once, promote the same artifact through environments. Rebuilding
  per environment means you're testing something different from what you ship.
- **Progressive delivery.** Small, frequent deploys behind feature flags, canary analysis, and
  automated rollback. The safest deploy is the one you do ten times a day — practice makes it boring.
- **The four golden signals.** Latency, traffic, errors, saturation — for every service. Alerts fire
  on symptoms (user-facing pain) with runbook links, not on every cause (a CPU spike nobody feels
  is not a 3am page).
- **SLOs and error budgets.** Define what "reliable enough" means numerically, then let the budget
  arbitrate the ship-vs-harden debate. No budget left? You fix reliability. Budget to spare? Ship.
- **Blameless postmortems.** Incidents are system failures, not people failures. The output is
  action items with owners and deadlines — "be more careful" is not an action item.

## Practical workflow

1. **Map the path to production.** Commit → build → test → artifact → stage → prod. Every manual
   step is a candidate for automation; every slow step gets measured.
2. **Codify infrastructure.** Start with the network and compute baseline in Terraform/Pulumi/
   CloudFormation; application config via environment, never baked into images; secrets via a
   secret manager, never in repos or env files committed anywhere.
3. **Build the deployment pipeline:** lint + unit → build immutable image (pinned digests) →
   integration tests → deploy to staging (production-like data shape) → smoke tests → progressive
   prod rollout (canary 1% → 25% → 100% with automatic rollback on SLO breach).
4. **Instrument before you need it.** Structured logs with correlation IDs, RED/USE metrics per
   service, distributed traces across boundaries, and dashboards for each service's golden signals.
5. **Define alerting tiers.** Page (user impact now, actionable, runbook attached) vs ticket
   (degraded, investigate soon) vs log (informational). Every page must be actionable — if the
   response is "wait and see," it's not a page.
6. **Practice incidents.** Game days / fire drills for the scary scenarios; postmortems within 48h
   of any significant incident; track action items to completion like product work.

Pipeline stage checklist:

```text
[ ] Build is hermetic and reproducible (pinned base images, locked deps)
[ ] Same artifact promoted across environments (no rebuilds)
[ ] Secrets injected at runtime from secret manager
[ ] Rollback is one command / automatic on health-check failure
[ ] Deploy is progressive (canary or blue/green), not big-bang
[ ] Post-deploy smoke tests verify user-critical paths
[ ] Dashboard + alerts exist before first prod traffic
```

## Common pitfalls

- **Snowflake environments.** "Works in staging" means nothing if staging differs from prod.
  Production-like data volume, config parity, and the same artifact — or your tests are theater.
- **Alert fatigue.** Paging on CPU thresholds and disk warnings trains on-call to ignore pages.
  Alert on user-impacting symptoms; tune ruthlessly; every false page is a bug in the alert.
- **Manual deploys with extra steps.** A wiki page titled "release process" with 30 manual steps is
  not a pipeline. If a human must remember the order, it will eventually go wrong at 2am.
- **Secrets in repos.** Even private repos leak (contractors, forks, history). Assume any committed
  secret is compromised: rotate and move to a manager.
- **No rollback plan.** "We'll roll forward" during an outage is optimism, not a plan. Fast,
  tested rollback beats heroic forward-fixing under pressure.
- **Monitoring everything, observing nothing.** 500 dashboards nobody opens. Start from questions
  ("is checkout healthy?") and build the minimum telemetry that answers them.
- **Treating toil as inevitable.** If on-call spends hours on repetitive tasks (log diving,
  manual scaling, cert renewals), that's engineering work waiting to be automated — budget for it.

Attribution

aicodedecodeaicodedecode
View sourceSee grades on GitHubMore from aicodedecode →
SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments (0)

No comments yet. Be the first to comment!

SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Related Skills

Caveman

Terse caveman voice: answer first, fluff gone, every technical fact kept. Use for /caveman, "caveman mode", "talk like caveman", "be brief", "less tokens". Stays on until "stop caveman" or "normal mode".

1100021 votes

Hyperplan

Adversarial multi-agent planning skill. Self-orchestrates 5 hostile category members (unspecified-low, unspecified-high, deep, ultrabrain, artistry) via team-mode for ruthless cross-critique debate, distills only the defensible insights, then MANDATORILY hands the distilled insight bundle to the `plan` agent for executable plan formalization. Use when planning needs maximum rigor and surfacing of weak assumptions, blind spots, and over-engineering. Triggers: 'hyperplan', 'hpp', '/hyperplan', ...

698621 votes

Writing Skills

Create and manage Claude Code skills in HASH repository following Anthropic best practices. Use when creating new skills, modifying skill-rules.json, understanding trigger patterns, working with hooks, debugging skill activation, or implementing progressive disclosure. Covers skill structure, YAML frontmatter, trigger types (keywords, intent patterns), UserPromptSubmit hook, and the 500-line rule. Includes validation and debugging with SKILL_DEBUG. Examples include rust-error-stack, cargo-dep...

3931 votes

Mcp Code Execution

Routes multi-tool workflows through MCP servers for large datasets and pipelines. Use when Bash tool overhead is limiting throughput on data-heavy tasks.

3421 votes

catchup

Recovers the conversation and failed tool calls of a previous Codex, Amp, Claude Code, Antigravity, Cline, Copilot CLI, Cursor, DeepSeek Harness, Grok Build, Kimi, OpenCode, Pi Agent, or ZCode session. Use when the user says "catch up", "what did the last session do", "get me up to speed", "I switched agents", asks to recover/summarize a previous session before continuing, or asks to diagnose or report a catchup failure. Do NOT use for the current conversation, git history, or any non-agent log.

741 votes
View all in ai-agents →