Skills DirectorySkills Directory
SkillsLearnSecurityCategoriesDocsBlogPro
Sign InSubmit Skill
Skills Directory

Security-tested agent skills for Claude, coding agents, and AI workflows.

Directory

  • Browse Skills
  • All Skills A–Z
  • Claude Skills
  • Claude Code Skills
  • Agent Skills
  • Categories
  • Authors
  • Submit a Skill

Learn

  • Learn Hub
  • Install Claude Skills
  • Write SKILL.md
  • Skills vs MCP
  • Directories Compared

Security

  • Security
  • Methodology
  • Secure Claude Skills
  • Security Badges
  • Chrome Extension
  • Skill Manager

Company

  • About
  • Community
  • Blog
  • API Docs
  • Advertise

2026 Skills Directory. All rights reserved.

ProTermsPrivacyRefunds
Back to skills

Aws Observability And Alerting

ASecurity

Step-by-step playbook for wiring CloudWatch metrics, alarms, dashboards, and X-Ray distributed tracing into an AWS workload — from log groups and metric filters to composite alarms and anomaly detection.

7 stars
0 votes
0 copies
0 views
Added 9/23/2026
ai-agentsexpressawsapidatabasesecurity

Works with

api

Security Analysis

A100/100

Scanned 9/23/2026

$npx -y skills add mcorbett51090/RavenClaude --skill aws-observability-and-alerting --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Aws Observability And Alerting?

Add the live security badge to your README — it updates automatically with every re-scan.

Security grade badge for Aws Observability And Alerting
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/mcorbett51090-aws-observability-and-alerting/badge)](https://www.skillsdirectory.com/skills/mcorbett51090-aws-observability-and-alerting)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
Files
SKILL.md
---
name: aws-observability-and-alerting
description: "Step-by-step playbook for wiring CloudWatch metrics, alarms, dashboards, and X-Ray distributed tracing into an AWS workload — from log groups and metric filters to composite alarms and anomaly detection."
---

# AWS Observability and Alerting

## When to invoke

Use when standing up or auditing the observability layer for a new or existing AWS workload: Lambda/ECS/EKS services, RDS, and API Gateway. Pair with `aws-finops` when cost anomaly alerting is in scope.

## Step 1 — Structured logging foundation

1. Every compute resource writes **JSON-structured logs** to a CloudWatch Log Group with a **retention policy** (never `never-expire`; 30/60/90 days for app logs, 365 days for security/audit).
2. Use **Embedded Metric Format (EMF)** to emit custom metrics from logs without a separate PutMetricData call — stamp `_aws.CloudWatchMetrics` in the log event.
3. Apply a **resource tag strategy** on log groups: `Environment`, `Service`, `Team` — these become the cost/attribution axis for log insights queries.

## Step 2 — Core metric alarms (minimum set)

| Signal | Metric | Threshold trigger |
|---|---|---|
| Lambda errors | `Errors` / `Throttles` per function | p-error-rate > 1% over 5 min |
| Lambda duration | `Duration` P99 | Approaching function timeout |
| ECS/Fargate CPU | `CPUUtilization` per service | > 80% sustained 5 min |
| RDS connections | `DatabaseConnections` | > 80% of `max_connections` |
| RDS freeable memory | `FreeableMemory` | < 200 MB |
| API Gateway 5xx | `5XXError` | Any spike > 5/min |
| SQS DLQ depth | `ApproximateNumberOfMessagesVisible` on DLQ | > 0 (any message) |

Use **Composite Alarms** to reduce alarm noise: group a service's error + latency + saturation alarms under one composite that fires the SNS/PagerDuty topic.

## Step 3 — Anomaly detection on variable signals

For metrics without a hard threshold (e.g., request rate that follows day-of-week curves):

```hcl
resource "aws_cloudwatch_metric_alarm" "request_rate_anomaly" {
  alarm_name          = "${var.service}-request-rate-anomaly"
  comparison_operator = "GreaterThanUpperThreshold"
  evaluation_periods  = 3
  threshold_metric_id = "e1"

  metric_query {
    id          = "m1"
    return_data = false
    metric { ... }
  }
  metric_query {
    id          = "e1"
    expression  = "ANOMALY_DETECTION_BAND(m1, 2)"
    return_data = true
  }
}
```

Band width `2` → ~95% confidence interval; tighten to `1.5` for latency-sensitive workloads.

## Step 4 — X-Ray distributed tracing

1. Enable **Active Tracing** on Lambda (environment variable `AWS_XRAY_SDK_ENABLED=true` or `TracingConfig: Active`).
2. For ECS: set the X-Ray daemon as a **sidecar container** with port 2000/UDP open in the task's security group.
3. Annotate segments with **service name, version, and customer/tenant id** — these become filterable in the X-Ray console and Service Map.
4. Set a **sampling rule**: 5% reservoir + 1/sec fixed rate for high-traffic paths; 100% for error paths (rule condition: `http.url` contains `/payment`).

## Step 5 — Dashboards

Build one CloudWatch Dashboard per service with:
- **Request rate + error rate + P50/P95/P99 latency** on row 1 (the RED signals).
- **Saturation signals** (CPU, memory, connections, queue depth) on row 2.
- **Downstream dependency health** (RDS, DynamoDB, external APIs via X-Ray) on row 3.
- Link the dashboard ARN in the service's runbook.

## Step 6 — Alarm routing

```
Alarm → SNS topic (per environment)
  ├── Sev-1: PagerDuty / OpsGenie (prod P99 latency, error rate, DLQ depth)
  └── Sev-2: Slack webhook (staging, cost anomalies, non-critical warnings)
```

Use **AWS Chatbot** to route CloudWatch alarms directly to Slack channels — avoids a Lambda forwarder for basic notification.

## Pitfalls

- **Setting `RetentionInDays: 0` (infinite) on log groups** — a common default that silently accumulates GBs and bill.
- **Missing DLQ depth alarms** — a silent DLQ is how async processing fails invisibly for days.
- **Composite alarms skipped** — individual alarms per metric create alert fatigue; the composite is the on-call signal, the individual alarms are the diagnostic detail.
- **X-Ray sampling at 100% in prod** — traces every request; at high throughput this adds cost and latency. Use reservoir + fixed-rate rules.
- **Dashboard with no link in the runbook** — a dashboard nobody finds during an incident has zero value.

Attribution

mcorbett51090mcorbett51090
View sourceSee grades on GitHubMore from mcorbett51090 →
SSkills Directory ProSkills Directory

Get any skill into Claude in one click.

Download any skill as a ZIP for Claude.ai, Claude Desktop, or .claude/skills. $9/mo.

See Pro

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments (0)

No comments yet. Be the first to comment!

SSkills Directory ProSkills Directory

Get any skill into Claude in one click.

Download any skill as a ZIP for Claude.ai, Claude Desktop, or .claude/skills. $9/mo.

See Pro

Related Skills

Caveman

Terse caveman voice: answer first, fluff gone, every technical fact kept. Use for /caveman, "caveman mode", "talk like caveman", "be brief", "less tokens". Stays on until "stop caveman" or "normal mode".

1100021 votes

Hyperplan

Adversarial multi-agent planning skill. Self-orchestrates 5 hostile category members (unspecified-low, unspecified-high, deep, ultrabrain, artistry) via team-mode for ruthless cross-critique debate, distills only the defensible insights, then MANDATORILY hands the distilled insight bundle to the `plan` agent for executable plan formalization. Use when planning needs maximum rigor and surfacing of weak assumptions, blind spots, and over-engineering. Triggers: 'hyperplan', 'hpp', '/hyperplan', ...

698431 votes

Writing Skills

Create and manage Claude Code skills in HASH repository following Anthropic best practices. Use when creating new skills, modifying skill-rules.json, understanding trigger patterns, working with hooks, debugging skill activation, or implementing progressive disclosure. Covers skill structure, YAML frontmatter, trigger types (keywords, intent patterns), UserPromptSubmit hook, and the 500-line rule. Includes validation and debugging with SKILL_DEBUG. Examples include rust-error-stack, cargo-dep...

3931 votes

Mcp Code Execution

Routes multi-tool workflows through MCP servers for large datasets and pipelines. Use when Bash tool overhead is limiting throughput on data-heavy tasks.

3421 votes

catchup

Recovers the conversation and failed tool calls of a previous Codex, Amp, Claude Code, Antigravity, Cline, Copilot CLI, Cursor, DeepSeek Harness, Grok Build, Kimi, OpenCode, Pi Agent, or ZCode session. Use when the user says "catch up", "what did the last session do", "get me up to speed", "I switched agents", asks to recover/summarize a previous session before continuing, or asks to diagnose or report a catchup failure. Do NOT use for the current conversation, git history, or any non-agent log.

741 votes
View all in ai-agents →