Skills DirectorySkills Directory
SkillsLearnSecurityCategoriesDocsBlogPro
Sign InSubmit Skill
Skills Directory

Security-tested agent skills for Claude, coding agents, and AI workflows.

Directory

  • Browse Skills
  • All Skills A–Z
  • Claude Skills
  • Claude Code Skills
  • Agent Skills
  • Categories
  • Authors
  • Submit a Skill

Learn

  • Learn Hub
  • Install Claude Skills
  • Write SKILL.md
  • Skills vs MCP
  • Directories Compared

Security

  • Security
  • Methodology
  • Secure Claude Skills
  • Security Badges
  • Chrome Extension
  • Skill Manager

Company

  • About
  • Community
  • Blog
  • API Docs
  • Advertise

2026 Skills Directory. All rights reserved.

ProTermsPrivacyRefunds
Back to skills

Prometheus Pro

ASecurity

Prometheus guidance — metrics design, PromQL, alerting rules, exporters, recording rules, and operations.

2 stars
0 votes
0 copies
0 views
Added 9/29/2026
ai-agentspythongonodekubernetesapi

Works with

cliapi

Security Analysis

A100/100

Scanned 9/29/2026

$npx -y skills add aicodedecode/awesome-muse-skills --skill prometheus-pro --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Prometheus Pro?

Add the live security badge to your README — it updates automatically with every re-scan.

Security grade badge for Prometheus Pro
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/aicodedecode-prometheus-pro/badge)](https://www.skillsdirectory.com/skills/aicodedecode-prometheus-pro)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
Files
SKILL.md
---
name: prometheus-pro
description: Prometheus guidance — metrics design, PromQL, alerting rules, exporters, recording rules, and operations.
category: development
---

## Overview

Prometheus is the standard open-source metrics system: it scrapes HTTP endpoints for time-series metrics, stores them efficiently, queries with PromQL, and alerts via Alertmanager. The pull model, dimensional labels, and powerful query language make it the backbone of Kubernetes observability and much more.

Using Prometheus well is about metric design (what to instrument, how to label), PromQL fluency, and alerting discipline (alert on symptoms, not causes). This skill covers instrumentation patterns, PromQL, recording/alerting rules, and operating Prometheus reliably.

## When to use

- Instrumenting applications (choosing metrics and labels).
- Writing PromQL queries and dashboards.
- Designing alerting rules (what to alert on, thresholds).
- Setting up exporters (node, blackbox, custom).
- Configuring recording rules for expensive queries.
- Operating Prometheus (retention, HA, federation, remote storage).
- Choosing between Prometheus, Datadog, and cloud metrics.

## Core concepts

- **Pull model.** Prometheus scrapes `/metrics` endpoints on an interval — targets can die without affecting the server, and there's no agent to deploy per app. Pushgateway exists for batch jobs; it's the exception, not the pattern.
- **Metric types.** Counter (monotonically increasing — requests, errors), Gauge (current value — connections, queue depth), Histogram (distributions — latencies, with buckets), Summary (client-computed quantiles — avoid, use histograms). Choosing the wrong type makes queries misleading.
- **Labels are dimensions.** `http_requests_total{method="GET", status="500"}` — labels slice metrics. Keep cardinality bounded: labels with unbounded values (user IDs, request IDs) explode storage and crash servers.
- **Naming conventions.** `<namespace>_<subsystem>_<unit>_{total,seconds,bytes}` — `http_request_duration_seconds`. Consistent names make PromQL discoverable; document your conventions.
- **RED and USE.** RED (Rate, Errors, Duration) for services; USE (Utilization, Saturation, Errors) for resources. These mnemonics structure both instrumentation and dashboards — cover them before exotic metrics.
- **PromQL essentials.** `rate()` for counters (always over a range, never `irate` for alerting), `histogram_quantile()` for latency percentiles, `avg_over_time`/`max_over_time` for gauges, label matching and aggregation (`sum by (service)`).
- **Recording rules.** Precompute expensive queries (`job:http_request_duration:p99`) on an interval — dashboards and alerts query the recording, not the raw data. Essential at scale.
- **Alerting rules.** Alert on symptoms (error rate, latency, saturation) with `for:` durations to avoid flapping. Every alert needs a runbook link, a severity, and an owner — alerts without runbooks are just noise.
- **Alertmanager.** Deduplication, grouping, silencing, and routing (PagerDuty/Opsgenie/Slack). Group related alerts; route by severity and team; inhibit cascading alerts (don't page for every pod when the node is down).
- **Exporters.** node_exporter (host metrics), blackbox_exporter (probing), kube-state-metrics (K8s objects), plus client libraries for app metrics. Instrument apps directly rather than parsing logs for metrics.
- **Service discovery.** Kubernetes, Consul, EC2, DNS-based discovery — targets come from the environment, not static configs. Relabeling shapes discovered targets into useful label sets.
- **Retention and storage.** Local TSDB with retention by time/size; 15 days default is often too short — size for your investigation needs. Remote write (Thanos, Cortex/Mimir, VictoriaMetrics) for long-term storage and HA.
- **HA.** Paired Prometheus servers (each scraping everything) + remote write, or Thanos/Cortex for global query. Single Prometheus is a single point of failure for alerting — plan accordingly.
- **Cardinality.** The #1 operational risk: unbounded label values → memory explosion → OOM. Monitor `prometheus_tsdb_head_series`, set label limits, and review new metrics' cardinality before deploying.
- **Exemplars.** Link metrics to traces (trace IDs on histograms) — connecting the RED metrics to the slow traces in one click.

## Practical workflow

1. **Instrument RED metrics.** Every service exposes request rate, error rate, and duration histogram with bounded labels (route templates, not raw paths).
   ```python
   REQUESTS = Counter("http_requests_total", "Requests", ["method", "route", "status"])
   LATENCY = Histogram("http_request_duration_seconds", "Latency", ["method", "route"])
   ```
2. **Deploy exporters.** node_exporter on hosts, kube-state-metrics in clusters, blackbox for external probing; app metrics via client libraries on `/metrics`.
3. **Write recording rules.** Precompute SLIs and expensive aggregations; name them clearly (`job:metric:aggregation`).
   ```yaml
   groups:
   - name: slis
     interval: 1m
     rules:
     - record: job:http_requests:error_ratio
       expr: sum by (job) (rate(http_requests_total{status=~"5.."}[5m]))
             / sum by (job) (rate(http_requests_total[5m]))
   ```
4. **Alert on symptoms.** Error-ratio, p99 latency, saturation alerts with `for:` and runbook annotations; severity tiers (page vs ticket).
   ```yaml
   - alert: HighErrorRate
     expr: job:http_requests:error_ratio > 0.05
     for: 5m
     labels: { severity: page }
     annotations: { runbook: https://wiki.example.com/runbooks/high-error-rate }
   ```
5. **Configure Alertmanager.** Grouping by alertname/cluster, routing by severity/team, inhibition for cascades, silences for maintenance — not for ignoring problems.
6. **Build dashboards from SLIs.** RED per service, USE per resource, recording-rule-backed panels; dashboards answer "is it broken and where" in under a minute.
7. **Manage cardinality.** Review label sets on new metrics; drop high-cardinality labels at scrape time via relabeling; alert on head-series growth.
8. **Plan retention and HA.** Size retention for investigation needs; remote-write to long-term storage; paired servers or Thanos for alerting HA.

## Common pitfalls

- **Unbounded label cardinality** — user IDs or request IDs as labels; OOM. Bound every label's value set.
- **Alerting on causes** — paging for CPU spikes instead of user-facing symptoms; alert on SLIs.
- **No `for:` duration** — flapping alerts on transient blips; require sustained breach.
- **Alerts without runbooks** — 3am pages with no guidance; every alert links a runbook.
- **`irate` in alerts** — spiky and misleading; `rate` over sensible windows.
- **Raw paths as labels** — `/users/123` exploding cardinality; use route templates.
- **Summary instead of histogram** — client-side quantiles can't aggregate; histograms with `histogram_quantile`.
- **Single Prometheus for alerting** — SPOF for pages; HA pair or remote-write + redundant alerting.
- **No recording rules at scale** — dashboards running expensive queries on every load; precompute.
- **Pushgateway as default** — for batch jobs only; services get scraped.
- **Ignoring silences hygiene** — permanent silences hiding real problems; expire and review them.
- **Retention too short** — 15 days when investigations need months; size deliberately with remote storage.
- **Scraping everything at 15s** — unnecessary load; tune intervals per target importance.

Attribution

aicodedecodeaicodedecode
View sourceSee grades on GitHubMore from aicodedecode →
SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments (0)

No comments yet. Be the first to comment!

SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Related Skills

Caveman

Terse caveman voice: answer first, fluff gone, every technical fact kept. Use for /caveman, "caveman mode", "talk like caveman", "be brief", "less tokens". Stays on until "stop caveman" or "normal mode".

1100021 votes

Hyperplan

Adversarial multi-agent planning skill. Self-orchestrates 5 hostile category members (unspecified-low, unspecified-high, deep, ultrabrain, artistry) via team-mode for ruthless cross-critique debate, distills only the defensible insights, then MANDATORILY hands the distilled insight bundle to the `plan` agent for executable plan formalization. Use when planning needs maximum rigor and surfacing of weak assumptions, blind spots, and over-engineering. Triggers: 'hyperplan', 'hpp', '/hyperplan', ...

698621 votes

Writing Skills

Create and manage Claude Code skills in HASH repository following Anthropic best practices. Use when creating new skills, modifying skill-rules.json, understanding trigger patterns, working with hooks, debugging skill activation, or implementing progressive disclosure. Covers skill structure, YAML frontmatter, trigger types (keywords, intent patterns), UserPromptSubmit hook, and the 500-line rule. Includes validation and debugging with SKILL_DEBUG. Examples include rust-error-stack, cargo-dep...

3931 votes

Mcp Code Execution

Routes multi-tool workflows through MCP servers for large datasets and pipelines. Use when Bash tool overhead is limiting throughput on data-heavy tasks.

3421 votes

catchup

Recovers the conversation and failed tool calls of a previous Codex, Amp, Claude Code, Antigravity, Cline, Copilot CLI, Cursor, DeepSeek Harness, Grok Build, Kimi, OpenCode, Pi Agent, or ZCode session. Use when the user says "catch up", "what did the last session do", "get me up to speed", "I switched agents", asks to recover/summarize a previous session before continuing, or asks to diagnose or report a catchup failure. Do NOT use for the current conversation, git history, or any non-agent log.

741 votes
View all in ai-agents →