Skills DirectorySkills Directory
SkillsLearnSecurityCategoriesDocsBlogPro
Sign InSubmit Skill
Skills Directory

Security-tested agent skills for Claude, coding agents, and AI workflows.

Directory

  • Browse Skills
  • All Skills A–Z
  • Claude Skills
  • Claude Code Skills
  • Agent Skills
  • Categories
  • Authors
  • Submit a Skill

Learn

  • Learn Hub
  • Install Claude Skills
  • Write SKILL.md
  • Skills vs MCP
  • Directories Compared

Security

  • Security
  • Methodology
  • Secure Claude Skills
  • Security Badges
  • Chrome Extension
  • Skill Manager

Company

  • About
  • Community
  • Blog
  • API Docs
  • Advertise

2026 Skills Directory. All rights reserved.

ProTermsPrivacyRefunds
Back to skills

Audit Observability

ASecurity

Identify missing logging, metrics, tracing, and alerting for production services.

10 stars
0 votes
0 copies
0 views
Added 10/6/2026
code-qualitypythongonodedebugginggitapidatabase

Works with

api

Security Analysis

A100/100

Scanned 10/6/2026

$npx -y skills add tomzx/agents --skill audit-observability --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Audit Observability?

Add the live security badge to your README — it updates automatically with every re-scan.

Security grade badge for Audit Observability
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/tomzx-audit-observability/badge)](https://www.skillsdirectory.com/skills/tomzx-audit-observability)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
Files
SKILL.md
---
name: audit-observability
description: Identify missing logging, metrics, tracing, and alerting for production services.
argument-hint: "[path or service-name]"
---

# Audit Observability

Scans the codebase and running services for missing or insufficient observability: logging, metrics, tracing, and alerting. Produces a gaps report ranked by risk so teams can prioritize adding instrumentation before issues arise in production.

## Prerequisites

- Working directory is the root of the repository
- Optional: `$1` — path or service name to scope the audit (defaults to `.`)
- Access to the codebase to search for instrumentation patterns
- Read `.sdlc/context/architecture.md` to understand the services and infrastructure

## Observability Pillars

| Pillar | Purpose | Examples |
|---|---|---|
| **Logging** | Record discrete events for debugging | Structured logs, error logs, audit logs |
| **Metrics** | Quantitative measurements over time | Request count, latency histogram, error rate |
| **Tracing** | Follow a request across service boundaries | Distributed trace IDs, span propagation |
| **Alerting** | Notify on-call when conditions degrade | PagerDuty, OpsGenie, Grafana alerts |

## Steps

1. Read `.sdlc/context/architecture.md` to identify services, endpoints, databases, and external dependencies.

2. Detect which observability libraries and frameworks are already in use:
   ```
   rg -n -i "(structlog|logging|logger|prometheus|datadog|opentelemetry|otel|jaeger|zipkin|sentry)" \
     -g '*.{py,ts,js,go}' ${1:-.} | head -40
   ```

3. Audit **logging** coverage:
   - Are errors logged with context (request ID, user ID, stack trace)?
   - Are external service calls logged (request/response, latency, status)?
   - Are business-critical operations logged (auth events, data mutations, payments)?
   - Is structured logging used consistently?
   - Are log levels appropriate (not everything at INFO/DEBUG)?
   ```
   rg -n -i "(except|catch|error|raise|throw)" \
     -g '*.{py,ts,js,go}' ${1:-.} | rg -v "log|logger|sentry|capture" | head -30
   ```
   Flag error handlers that catch exceptions without logging them.

4. Audit **metrics** coverage:
   - Are HTTP endpoints instrumented (request count, latency, error rate)?
   - Are database queries instrumented (query latency, connection pool usage)?
   - Are external service calls instrumented (call count, latency, error rate)?
   - Are business metrics tracked (orders placed, emails sent, jobs completed)?
   - Are resource metrics available (CPU, memory, disk, connections)?
   ```
   rg -n -i "(counter|histogram|gauge|summary|metric|observe|inc\(|time\()" \
     -g '*.{py,ts,js,go}' ${1:-.} | head -30
   ```
   Identify endpoints and services with no metrics instrumentation.

5. Audit **tracing** coverage:
   - Is OpenTelemetry or equivalent configured?
   - Are trace IDs propagated across service boundaries (HTTP headers, message queues)?
   - Are database calls and external HTTP calls included in traces?
   - Are spans created for significant operations?
   ```
   rg -n -i "(trace_id|span_id|opentelemetry|otel|tracer|start_span|propagat)" \
     -g '*.{py,ts,js,go}' ${1:-.} | head -30
   ```

6. Audit **alerting** coverage:
   - Are alerts defined for SLO breaches (error rate, latency)?
   - Are alerts defined for infrastructure issues (disk full, memory pressure, connection exhaustion)?
   - Are alerts defined for business anomalies (queue depth, failed jobs, payment failures)?
   - Are alert thresholds appropriate (not too noisy, not too quiet)?
   - Look for alert configuration files:
   ```
   find ${1:-.} -type f \( -name "alerts*" -o -name "alerting*" -o -name "slo*" -o -name "rules*" \) \
     | rg -v node_modules | rg -v .git | head -20
   ```

7. Rank gaps by risk:
   - **High risk:** No error logging on critical paths, no alerts for SLO breaches, no metrics on user-facing endpoints
   - **Medium risk:** Missing tracing on cross-service calls, no metrics on background jobs, inconsistent log levels
   - **Low risk:** Missing business metrics, non-critical paths without tracing, verbose logging in non-critical paths

8. Produce the gaps report.

## Output Format

```markdown
# Observability Audit — <service-name or project>

**Date:** <YYYY-MM-DD>
**Scope:** <path or services audited>

## Summary

- **Logging:** Adequate / Partial / Missing
- **Metrics:** Adequate / Partial / Missing
- **Tracing:** Adequate / Partial / Missing
- **Alerting:** Adequate / Partial / Missing
- **Tools detected:** <prometheus, structlog, opentelemetry, sentry, etc.>
- **High-risk gaps:** N
- **Medium-risk gaps:** N
- **Low-risk gaps:** N

## Logging Gaps

| Risk | Location | Issue | Recommendation |
|---|---|---|---|
| High | `<file>:<line>` | Error swallowed without logging | Add structured error log with context |
| Medium | `<file>:<line>` | External call not logged | Add request/response logging |

## Metrics Gaps

| Risk | Endpoint / Operation | Issue | Recommendation |
|---|---|---|---|
| High | `POST /api/checkout` | No latency or error metrics | Add histogram for latency, counter for errors |
| Medium | Background job processor | No job count or duration metrics | Add counter and histogram |

## Tracing Gaps

| Risk | Boundary | Issue | Recommendation |
|---|---|---|---|
| High | Service A → Service B | Trace ID not propagated | Add trace context headers |
| Medium | Database queries | No span for DB calls | Add DB query spans |

## Alerting Gaps

| Risk | Condition | Issue | Recommendation |
|---|---|---|---|
| High | Error rate > 1% | No alert configured | Add SLO-based alert |
| Medium | Queue depth > 1000 | No alert configured | Add threshold alert |

## Recommended Implementation Order

1. <Highest-priority gap to fix first>
2. <Next priority>
3. ...
```

## Example Usage

**Scenario 1: New project review**
```
/audit-observability
```
Project has logging via structlog but no metrics, tracing, or alerting. Recommends adding Prometheus for metrics and OpenTelemetry for tracing as the top priorities.

**Scenario 2: Pre-production readiness**
```
/audit-observability src/api
```
Scanning the API layer before going to production. Finds 3 endpoints without latency metrics and 2 error handlers that catch exceptions without logging. Recommends adding metrics and error logging before launch.

**Scenario 3: Post-incident follow-up**
```
/audit-observability
```
After an incident caused by undetected error rate spike. Finds no alerts configured for 5xx rate. Recommends adding SLO-based alerting as the top priority.

## Next Step

If gaps are found, create issues for the highest-priority items via `/create-issue`.
To check current production health, run `/observe-production`.

## Useful Commands Reference

| Command | Description |
|---|---|
| `rg -n "structlog\|logging\|logger" -g '*.py' . | head -30` | Find logging usage in Python |
| `rg -n "prometheus\|metrics\|Counter\|Histogram" -g '*.py' . | head -30` | Find metrics instrumentation |
| `rg -n "opentelemetry\|otel\|trace" -g '*.py' . | head -30` | Find tracing setup |
| `find . -name "alerts*" -o -name "slo*" -o -name "rules*"` | Find alert configuration files |

Attribution

tomzxtomzx
View sourceSee grades on GitHubMore from tomzx →
SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments (0)

No comments yet. Be the first to comment!

SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Related Skills

Caveman Review

Ultra-compressed code review comments. Cuts noise from PR feedback while preserving the actionable signal. Each comment is one line: location, problem, fix. Use when user says "review this PR", "code review", "review the diff", "/review", or invokes /caveman-review. Auto-triggers when reviewing pull requests.

1100021 votes

Caveman Commit

Ultra-compressed commit message generator. Cuts noise from commit messages while preserving intent and reasoning. Conventional Commits format. Subject ≤50 chars, body only when "why" isn't obvious. Use when user says "write a commit", "commit message", "generate commit", "/commit", or invokes /caveman-commit. Auto-triggers when staging changes.

1100021 votes

Springboot Verification

Verification loop for Spring Boot projects: build, static analysis, tests with coverage, security scans, and diff review before release or PR.

2456590 votes

Verification Loop

一个全面的 Claude Code 会话验证系统。

2456590 votes

Django Verification

Verification loop for Django projects: migrations, linting, tests with coverage, security scans, and deployment readiness checks before release or PR.

2456590 votes
View all in code-quality →