Skills DirectorySkills Directory
SkillsLearnSecurityCategoriesDocsBlogPro
Sign InSubmit Skill
Skills Directory

Security-tested agent skills for Claude, coding agents, and AI workflows.

Directory

  • Browse Skills
  • All Skills A–Z
  • Claude Skills
  • Claude Code Skills
  • Agent Skills
  • Categories
  • Authors
  • Submit a Skill

Learn

  • Learn Hub
  • Install Claude Skills
  • Write SKILL.md
  • Skills vs MCP
  • Directories Compared

Security

  • Security
  • Methodology
  • Secure Claude Skills
  • Security Badges
  • Chrome Extension
  • Skill Manager

Company

  • About
  • Community
  • Blog
  • API Docs
  • Advertise

2026 Skills Directory. All rights reserved.

ProTermsPrivacyRefunds
Back to skills

Sre Incident Responder

ASecurity

当线上服务故障/告警触发、需要快速止损与事后复盘时使用;做事件指挥、分级定级、稳态恢复与无指责复盘(产出事件时间线、状态更新、复盘报告与改进项);不适用于纯功能开发、单元测试或非线上的日常运维;触发词:线上故障、事件响应、止损、复盘、SLO 燃尽

3 stars
0 votes
0 copies
2 views
Added 9/19/2026
ai-agentsdevops

Works with

cursorcli

Security Analysis

A100/100

Scanned 9/19/2026

$npx -y skills add findscripter/everything-skills --skill sre-incident-responder --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Sre Incident Responder?

Add the live security badge to your README — it updates automatically with every re-scan.

Security grade badge for Sre Incident Responder
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/findscripter-sre-incident-responder/badge)](https://www.skillsdirectory.com/skills/findscripter-sre-incident-responder)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
Files
SKILL.md
---
name: sre-incident-responder
title: SRE 事件响应
description: 当线上服务故障/告警触发、需要快速止损与事后复盘时使用;做事件指挥、分级定级、稳态恢复与无指责复盘(产出事件时间线、状态更新、复盘报告与改进项);不适用于纯功能开发、单元测试或非线上的日常运维;触发词:线上故障、事件响应、止损、复盘、SLO 燃尽
domain: 研发/devops
triggers: [线上故障, P0, 事件响应, 止损, 复盘, post-mortem, SLO 燃尽, 故障定级, on-call, 回滚]
tags: [sre, 事件响应, 可观测性, 故障定级, 复盘, 止损, on-call]
level: 进阶
status: stable
agents: [claude-code, codex, cursor, gemini-cli]
tools: [Prometheus, Grafana, OpenTelemetry, PagerDuty, Slack]
requires: []
related: []
combines_with: []
license: MIT
source: sickn33/agentic-awesome-skills
source_license: MIT
---
## 何时使用

适用:
- 线上服务发生中断、降级或大面积告警,需要在分钟级内组织响应、快速止损。
- 需要对事件分级(P0–P3 / SEV-1–4)、设立事件指挥、对内对外同步进度。
- 事件恢复后要做无指责复盘、根因分析并落地改进项。

不该用(负边界):
- 纯功能开发、写单测、代码评审等与线上故障无关的任务。
- 非紧急的日常运维变更、容量规划(除非由事件触发)。
- 缺少访问权限、监控数据或决策授权时——先停下来索要必要输入与边界,再行动。

核心心法:止损优先于追根因,准确高于速度(错误修复会指数级放大故障);持续、按受众深度同步;全程留痕。

## 步骤

1. 评估影响与定级(前 5 分钟)
   - 用户影响:受影响用户数、地域、关键链路。
   - 业务影响:营收损失、SLA 违约、体验劣化。
   - 系统范围:受影响服务、依赖、爆炸半径(blast radius)。
   - 据此定级(见下「严重度分级」)。

2. 建立事件指挥(Incident Command)
   - Incident Commander:唯一决策人,统筹响应。
   - Communication Lead:负责干系人与对外同步。
   - Technical Lead:统筹技术排查与修复。
   - 拉起作战室(频道 / 视频 / 共享文档)。

3. 立即稳态化(快速止损手段优先)
   - 快速止损:限流、特性开关(feature flag)、熔断。
   - 回滚评估:近期发布、配置变更、基础设施变更。
   - 扩容:自动/手动扩容、流量重分布。
   - 首条状态页与内部通知。

4. 可观测性驱动排查
   - 分布式追踪:OpenTelemetry / Jaeger / Zipkin 看请求链路。
   - 指标关联:Prometheus / Grafana / DataDog 找异常模式。
   - 日志聚合:ELK / Splunk / Loki 分析错误模式。
   - SRE 技法:SLI/SLO 违约与燃尽率(burn rate)、变更关联、依赖映射、级联失败(熔断状态、重试风暴、惊群)、容量与配额耗尽分析。

5. 修复与恢复
   - 最小可行修复(MVF):最快恢复路径。
   - 风险评估 + 灰度发布 + 校验 + 加强监控。
   - 恢复校验:所有 SLI 回到阈值内、真实用户监控正常、依赖健康、留有容量余量。

6. 事后流程
   - 24h 内:持续监控、宣布恢复、导出数据、团队 debrief。
   - 无指责复盘:时间线、根因分析(5 Whys / 鱼骨图 / 系统思维)、贡献因素(人/流程/技术债)、可跟踪的改进项。
   - 系统改进:监控与告警、自动化/自愈、韧性架构、流程与培训。

## 指令

- 先澄清目标、约束与所需输入;缺权限/数据/成功标准就停下来问。
- 全程对内每 15 分钟同步一次(活跃事件);对外维护状态页并给客服话术。
- 文档化:时间线(带时间戳)、决策理由、影响指标、沟通记录。
- 优先恢复服务,根因分析放到恢复之后。

严重度分级(保留源约束):
- P0 / SEV-1 完全中断或安全入侵:7x24 立即升级;确认 < 15 分钟,恢复 < 1 小时;每 15 分钟同步并通知高管。
- P1 / SEV-2 核心功能严重降级:确认 < 1 小时,恢复 < 4 小时;每小时同步并更新状态页。
- P2 / SEV-3 次要功能受影响:确认 < 4 小时,恢复 < 24 小时;按需内部同步。
- P3 / SEV-4 外观问题、无用户影响:下个工作日处理,恢复 < 72 小时;走标准工单。

韧性模式:熔断器、舱壁隔离(bulkhead)、优雅降级、重试策略(指数退避 + 抖动 jitter + 熔断)。
关键指标:MTTR、MTTD、事件频次、用户影响。

## 示例

场景:支付服务 P0 全站超时。
1. 定级 P0,拉作战室,指定 IC / Comms / Tech Lead;发首条状态页。
2. 查 Grafana 发现错误率在某次发布后陡增 → 用变更关联锁定可疑发布。
3. 立即止损:对该服务回滚上一个版本 + 限流保护下游;错误率回落。
4. 校验所有 SLI 回到阈值内、真实用户监控恢复、依赖健康,确认稳定。
5. 宣布恢复,24h 内出无指责复盘:根因为发布未覆盖连接池配置,改进项为「上线前连接池压测 + 自动回滚阈值」并跟踪闭环。

## 注意事项

- 速度重要,但正确更重要:错误修复会让局势指数级恶化。
- 止损优先,追根因靠后;活跃事件期不纠缠根因。
- 无指责文化:聚焦系统与流程,而非追责个人,保障心理安全。
- 数据驱动决策:基于可观测性与指标,而非猜测。
- 本技能产出不能替代针对具体环境的验证、测试与专家评审。

## 互见

- 配套平台:PagerDuty / Opsgenie(告警与排班)、ServiceNow(ITSM 与变更关联)、Slack/Teams(ChatOps 与自动同步)。
- 可观测性集成:统一仪表盘、告警关联降噪、Runbook 自动化诊断、事件回放。

---
采编自 sickn33/antigravity-awesome-skills(MIT 许可)。

Attribution

findscripterfindscripter
View sourceSee grades on GitHubMore from findscripter →
SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments (0)

No comments yet. Be the first to comment!

SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Related Skills

Caveman

Terse caveman voice: answer first, fluff gone, every technical fact kept. Use for /caveman, "caveman mode", "talk like caveman", "be brief", "less tokens". Stays on until "stop caveman" or "normal mode".

1100021 votes

Hyperplan

Adversarial multi-agent planning skill. Self-orchestrates 5 hostile category members (unspecified-low, unspecified-high, deep, ultrabrain, artistry) via team-mode for ruthless cross-critique debate, distills only the defensible insights, then MANDATORILY hands the distilled insight bundle to the `plan` agent for executable plan formalization. Use when planning needs maximum rigor and surfacing of weak assumptions, blind spots, and over-engineering. Triggers: 'hyperplan', 'hpp', '/hyperplan', ...

698621 votes

Writing Skills

Create and manage Claude Code skills in HASH repository following Anthropic best practices. Use when creating new skills, modifying skill-rules.json, understanding trigger patterns, working with hooks, debugging skill activation, or implementing progressive disclosure. Covers skill structure, YAML frontmatter, trigger types (keywords, intent patterns), UserPromptSubmit hook, and the 500-line rule. Includes validation and debugging with SKILL_DEBUG. Examples include rust-error-stack, cargo-dep...

3931 votes

Mcp Code Execution

Routes multi-tool workflows through MCP servers for large datasets and pipelines. Use when Bash tool overhead is limiting throughput on data-heavy tasks.

3421 votes

catchup

Recovers the conversation and failed tool calls of a previous Codex, Amp, Claude Code, Antigravity, Cline, Copilot CLI, Cursor, DeepSeek Harness, Grok Build, Kimi, OpenCode, Pi Agent, or ZCode session. Use when the user says "catch up", "what did the last session do", "get me up to speed", "I switched agents", asks to recover/summarize a previous session before continuing, or asks to diagnose or report a catchup failure. Do NOT use for the current conversation, git history, or any non-agent log.

741 votes
View all in ai-agents →