Skills DirectorySkills Directory
SkillsLearnSecurityCategoriesDocsBlogPro
Sign InSubmit Skill
Skills Directory

Security-tested agent skills for Claude, coding agents, and AI workflows.

Directory

  • Browse Skills
  • All Skills A–Z
  • Claude Skills
  • Claude Code Skills
  • Agent Skills
  • Categories
  • Authors
  • Submit a Skill

Learn

  • Learn Hub
  • Install Claude Skills
  • Write SKILL.md
  • Skills vs MCP
  • Directories Compared

Security

  • Security
  • Methodology
  • Secure Claude Skills
  • Security Badges
  • Chrome Extension
  • Skill Manager

Company

  • About
  • Community
  • Blog
  • API Docs
  • Advertise

2026 Skills Directory. All rights reserved.

ProTermsPrivacyRefunds
Back to skills

Create Observability

ASecurity

Define logging, metrics, tracing, and alerting for a feature so production health can be monitored from day one.

10 stars
0 votes
0 copies
0 views
Added 10/6/2026
devopsgodebuggingapi

Works with

api

Security Analysis

A100/100

Scanned 10/6/2026

$npx -y skills add tomzx/agents --skill create-observability --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Create Observability?

Add the live security badge to your README — it updates automatically with every re-scan.

Security grade badge for Create Observability
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/tomzx-create-observability/badge)](https://www.skillsdirectory.com/skills/tomzx-create-observability)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
Files
SKILL.md
---
name: create-observability
description: Define logging, metrics, tracing, and alerting for a feature so production health can be monitored from day one.
argument-hint: "[specification-doc]"
---

# Create Observability

Defines how a feature's production health will be monitored by identifying log statements, service metrics, distributed traces, and alerts before implementation begins.

Without this step, features go to production with no monitoring: outages go undetected, root causes take hours to find, and on-call engineers lack runbooks.

## Prerequisites

- Apply the shared SDLC conventions in `skills/sdlc/references/shared.md`.
- If no argument is provided, locate the feature directory under `.sdlc/features/` whose frontmatter `issue` field references `$ISSUE_NUMBER`.
- `.sdlc/features/N-<slug>/specification.md` (must have passed review with findings verdict `approved`), or a specification document provided in context or as a file path (`$1`)
- `.sdlc/features/N-<slug>/lifecycle.md` (optional, if a lifecycle document was produced): monitor state transitions, invariant violations, and retention cleanup as part of observability
- `.sdlc/features/N-<slug>/telemetry.md` (optional, if a telemetry plan was produced): align observability with business metrics already defined
- `.sdlc/features/N-<slug>/requirements.md` (optional, for cross-referencing NFRs like latency and availability targets)

## Steps

1. Read the specification, lifecycle document (if present), telemetry plan (if present), and requirements.
2. Identify the critical paths and failure modes from the specification.
3. For each critical path, determine what log entries are needed for debugging.
4. Define service-level metrics (counters, histograms, gauges) that reflect system health.
5. Identify where distributed traces should be emitted for cross-service flows.
6. Define health checks and readiness probes for new services or endpoints.
7. Specify alerts with clear conditions, severity, and runbook links. When the monitoring stack is Prometheus-compatible, write the normative alert definitions to `.sdlc/features/N-<slug>/alerts.yaml` (Prometheus rule format, template at `skills/sdlc/templates/features/alerts.yaml`) and keep the per-alert tables in the document as the human-readable summary. Validate best-effort with `promtool check rules alerts.yaml` when available; a missing tool is skipped, a validation failure is a defect to fix before handoff.
8. Determine observability infrastructure requirements (existing vs. new instrumentation).
9. Write the output to `.sdlc/features/N-<slug>/observability.md` (plus `alerts.yaml` when alerts are defined and the stack is Prometheus-compatible).

## Output Format

Use the template at `skills/sdlc/templates/features/observability.md` (copied to `.sdlc/templates/features/observability.md` by `/initialize-sdlc-directory`; use the project's customized copy if present). Write the result to the artifact path named in the steps above.

## Logging Guidance

- Use structured logging (JSON or key-value) so logs are queryable.
- Log at the boundary of the system (incoming requests, outgoing calls to external services, state transitions).
- Include a `correlation_id` or `trace_id` on every log entry to enable cross-service debugging.
- Avoid logging sensitive data (PII, secrets, tokens).
- Log levels: `DEBUG` (development only), `INFO` (normal operations), `WARN` (degraded but recoverable), `ERROR` (unexpected failure requiring attention).

## Metrics Guidance

Common metric types for features:

- **Request rate:** Number of requests per second to new endpoints.
- **Error rate:** Percentage of requests resulting in errors (4xx/5xx).
- **Latency:** Histogram of request durations (p50, p95, p99).
- **Queue depth:** Number of items pending processing (for async features).
- **Resource utilization:** CPU, memory, connections used by the feature.
- **Business metrics:** Counts tied to domain events (orders placed, files uploaded).

Every metric should answer: "If this number changes unexpectedly, what action do I take?" If no action exists, the metric is noise.

## Alert Guidance

Good alerts are:
- **Actionable:** Every alert triggers a human response. If nobody acts, remove the alert.
- **Specific:** The condition clearly identifies what is wrong, not just "something is slow."
- **Timely:** Fires fast enough to mitigate impact, but with a `for` duration to avoid flapping.
- **Sized correctly:** Critical alerts wake someone up; Warning alerts appear in dashboards; Info alerts are logged.

Every alert must have a runbook: a short list of steps to diagnose and resolve.

## Tracing Guidance

Add spans at service boundaries and for expensive operations (DB queries, external API calls, large computations). Record attributes that help narrow down the issue: user ID, request ID, operation type, resource identifier.

## Outcome

If `$OUTCOME_YAML` is set, emit `verdict: approved` there per `skills/sdlc/references/shared.md`, If the artifact could not be produced, omit the file.
In the same emission, list every file you produced under `artifacts:` (`.sdlc/features/N-<slug>/observability.md`, plus `.sdlc/features/N-<slug>/alerts.yaml` when written).

## Example Usage

**Scenario 1: REST API endpoint**
Specification defines `POST /orders`.
Metrics: `orders_request_total` (counter), `orders_request_duration_seconds` (histogram), `orders_error_total` (counter by status code).
Logging: INFO on order created (with order_id, user_id), WARN on validation failure, ERROR on DB write failure.
Alert: fire Critical if error rate > 5% for 5 minutes. Runbook: check DB connectivity, check upstream service.
SLO: 99.9% availability, p99 latency < 500ms.

**Scenario 2: Background job**
Specification defines a nightly data export.
Metrics: `export_jobs_total` (counter), `export_duration_seconds` (histogram), `export_records_processed` (counter).
Logging: INFO on job start/complete (with job_id, record_count), WARN on partial failure, ERROR on full failure.
Tracing: root span for the job, child spans for each batch.
Alert: fire Warning if job duration exceeds 2x normal, Critical if job fails 2 consecutive runs.

**Scenario 3: WebSocket connection**
Specification defines a real-time notification feed.
Metrics: `ws_connections_active` (gauge), `ws_messages_sent_total` (counter), `ws_connection_duration_seconds` (histogram).
Logging: INFO on connect/disconnect (with user_id), WARN on reconnect storm, ERROR on message delivery failure.
Health check: readiness probe that verifies the WebSocket server can accept connections.

## Completion Checklist

Before handing off to review, confirm:

- [ ] Each alert is actionable, severity-tagged, and links a runbook
- [ ] `alerts.yaml` written when the stack is Prometheus-compatible, and it agrees with the alert summary tables
- [ ] SLO/error-budget targets referenced from requirements where applicable

Self-check the draft against the [`review-observability` checklist](../review-observability/SKILL.md) and fix what you can, so review finds less to flag.

## Next Step

A review subagent is dispatched automatically to run `/review-observability` to audit the observability plan for completeness, actionability, and consistency before moving on.
Once approved, continue with `/create-plan`.

## Useful Commands Reference

| Command | Description |
|---|---|
| `promtool check rules alerts.yaml` | Best-effort Prometheus rule validation (skip and note when unavailable) |

Attribution

tomzxtomzx
View sourceSee grades on GitHubMore from tomzx →
SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments (0)

No comments yet. Be the first to comment!

SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Related Skills

Terraform Module Library

Build reusable Terraform modules for AWS, Azure, and GCP infrastructure following infrastructure-as-code best practices. Use when creating infrastructure modules, standardizing cloud provisioning, or implementing reusable IaC components.

401991 votes

sematext-otel

Wire a service's OpenTelemetry output to Sematext Cloud. Walks through region, App-type, instrumentation flow (managed OTLP endpoint vs Sematext Agent), and signal selection (traces/metrics/logs), then produces the exact env-var block and points at a runnable reference example in this repo. Invoke when instrumenting a new app for Sematext.

01 votes

Deployment Patterns

Deployment workflows, CI/CD pipeline patterns, Docker containerization, health checks, rollback strategies, and production readiness checklists for web applications. Use when setting up deployment infrastructure or planning releases.

2699140 votes

Babysit

Watch a pull request or review cycle until it is ready to merge. Use when asked to babysit, monitor, or keep checking PR comments, reviews, and CI until all actionable issues are resolved.

971540 votes

V7 Roster

Interact with the Paperclip control plane API for task coordination and governance. Use when checking assignments, updating issue status, posting comments, delegating work, managing routines, or calling Paperclip API endpoints.

953190 votes
View all in devops →