Skip to content
Back to skills

Incident Response

ASecurity

How to work a production incident remediation task. Use when the task title starts with "Incident:" or its description carries an incident id and a suggest/auto_fix policy.

  • 109 stars
  • 0 votes
  • 0 copies
  • 2 views
  • Added September 19, 2026
ai-agentsgodatabase

Security analysis

A100/100

Scanned October 2, 2026

npx -y skills add makifbaysal/tasktrooper --skill incident-response --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Incident Response?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Incident Response
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/makifbaysal-incident-response/badge)](https://www.skillsdirectory.com/skills/makifbaysal-incident-response)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: incident-response
category: operations
description: How to work a production incident remediation task. Use when the task title starts with "Incident:" or its description carries an incident id and a suggest/auto_fix policy.
---
# Incident response

A task titled `Incident: …` is a live production problem, assigned to you by the system. Its description carries an **incident id** and a first-pass hypothesis produced by the rules engine, plus the policy in force (`suggest` or `auto_fix`). That hypothesis is a starting point, not a verdict.

## Order of work

1. **Read the incident first.** `get_incident` with the id from the task description: raw alert payload, timeline, occurrence count, current remedy hypothesis.
2. **Correlate with the tools you hold.** `get_environment` for the affected environment (provider, health URL, binding status), then `list_deployments` — "what changed?" answers most incidents. `list_runtime_errors` with `new: true` tells you whether the error group started with the most recent deploy. `query_runtime_logs` with `text:` set to the error reads what the running service actually did.
3. **Classify.** Config/credentials, dependency, capacity, or code defect. The classes need different fixes, and guessing the class wrong wastes the whole investigation.
4. **Write the remedy.** `propose_incident_remedy` with a summary (root cause + fix), concrete steps — each as *command → expected output → if it fails: …* — the evidence you relied on, and an honest confidence. Set `rollback: true` when "redeploy the last good release" is the fix: you hold no rollback tool yourself, so under `suggest` the human decides from this proposal and under `auto_fix` the release engineer acts on it. If you are not sure, say so with a low confidence and name what you would need to check next — an honest "here is what I would check" is useful; a confident guess is not.
5. **Act per policy.**
   - **suggest** (default): do NOT change production code. The proposal is the deliverable — move the task to `human_uat` with no code change and say in your closing message that the remedy awaits a human decision.
   - **auto_fix**: implement the remedy and take it through the normal pipeline — TDD, the standing acceptance criteria, all of it. A hotfix that skips them causes the next incident.
6. **Resolve only after verifying recovery.** `resolve_incident` only after `get_environment` / `list_runtime_errors` show the environment is actually healthy again, with the note naming that evidence. Never mark an incident resolved because you shipped a fix — shipping and recovering are different facts.

## Rules

- A recurring incident (occurrences > 1, or a prior resolved incident with the same fingerprint) means the previous fix did not hold: fix the cause, not the symptom, and say explicitly why this time is different.
- Never write a remedy you cannot execute step by step. "Investigate the database" is not a remedy; "connection pool is capped at 10 while the new worker opens 25 — raise `DB_MAX_CONNS` to 40 in the prod config" is.
- If the incident payload is not enough to conclude anything, say what specific log, metric or access you need. Do not invent a cause to have something to write.

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…