Skills DirectorySkills Directory
SkillsLearnSecurityCategoriesDocsBlogPro
Sign InSubmit Skill
Skills Directory

Security-tested agent skills for Claude, coding agents, and AI workflows.

Directory

  • Browse Skills
  • All Skills A–Z
  • Claude Skills
  • Claude Code Skills
  • Agent Skills
  • Categories
  • Authors
  • Submit a Skill

Learn

  • Learn Hub
  • Install Claude Skills
  • Write SKILL.md
  • Skills vs MCP
  • Directories Compared

Security

  • Security
  • Methodology
  • Secure Claude Skills
  • Security Badges
  • Chrome Extension
  • Skill Manager

Company

  • About
  • Community
  • Blog
  • API Docs
  • Advertise

2026 Skills Directory. All rights reserved.

ProTermsPrivacyRefunds
Back to skills

Prometheus

ASecurity

Expert agent for Prometheus 3.x covering PromQL, scrape configuration, service discovery, recording rules, alerting rules, Alertmanager, TSDB management, native histograms, cardinality management, federation, remote write, and Thanos/Mimir integration. WHEN: \"Prometheus\", \"PromQL\", \"scrape config\", \"Alertmanager\", \"recording rule\", \"alerting rule\", \"TSDB\", \"remote write\", \"Thanos\", \"Mimir\", \"Pushgateway\", \"node_exporter\", \"histogram_quantile\", \"rate()\", \"relabel_c...

4 stars
0 votes
0 copies
0 views
Added 9/24/2026
devopsgonodeexpresskubernetesapidatabaseperformance

Works with

cliapi

Security Analysis

A100/100

Pro scans all 9 files and shows the line behind each finding

Scanned 9/24/2026

$npx -y skills add chrishuffman5/domain-expert --skill prometheus --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Prometheus?

Add the live security badge to your README — it updates automatically with every re-scan.

Security grade badge for Prometheus
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/chrishuffman5-prometheus/badge)](https://www.skillsdirectory.com/skills/chrishuffman5-prometheus)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
Files
SKILL.md
---
name: prometheus
description: "Expert agent for Prometheus 3.x covering PromQL, scrape configuration, service discovery, recording rules, alerting rules, Alertmanager, TSDB management, native histograms, cardinality management, federation, remote write, and Thanos/Mimir integration. WHEN: \"Prometheus\", \"PromQL\", \"scrape config\", \"Alertmanager\", \"recording rule\", \"alerting rule\", \"TSDB\", \"remote write\", \"Thanos\", \"Mimir\", \"Pushgateway\", \"node_exporter\", \"histogram_quantile\", \"rate()\", \"relabel_configs\", \"service discovery Prometheus\", \"native histogram\", \"exemplar\", \"cardinality\"."
license: MIT
---

# Prometheus Technology Expert

This skill covers Prometheus 3.x: PromQL, scrape configuration, service discovery, alerting, TSDB internals, and long-term storage integration. Every recommendation addresses the tradeoff triangle: **query performance**, **storage cost**, and **cardinality management**.

## How to Approach Tasks

When you receive a request:

1. **Classify** the request by area:
   - **Architecture** (pull model, TSDB, service discovery, remote write, HA) -- Load `references/architecture.md`
   - **PromQL** (queries, functions, aggregations, binary ops, examples) -- Load `references/promql.md`
   - **Alerting** (alert rules, Alertmanager, routing, receivers, fatigue) -- Load `references/alerting.md`
   - **Configuration** (prometheus.yml, scrape configs, relabeling, recording rules) -- Load `references/configuration.md`
   - **Diagnostics** (TSDB status, cardinality, scrape failures, memory, WAL) -- Load `references/diagnostics.md`

2. **Include cardinality context** -- Never recommend a metric or label without considering its cardinality impact. High cardinality is the primary cause of Prometheus performance issues.

3. **Recommend recording rules** -- For any expensive PromQL expression used in dashboards or alerts, suggest pre-computing it as a recording rule.

4. **Default to native histograms** -- Prometheus 3.x native histograms are more accurate and lower-cardinality than classic histograms. Recommend them for new instrumentation.

5. **Address long-term storage** -- Prometheus local TSDB is designed for short-term retention (15-30 days). For longer retention, recommend remote_write to Thanos, Mimir, or VictoriaMetrics.

## Core Expertise

- **Architecture:** Pull-based scraping, TSDB internals (WAL, blocks, compaction), service discovery (Kubernetes, Consul, EC2, file-based, DNS, HTTP), federation, remote read/write, Pushgateway for short-lived jobs
- **PromQL:** Instant and range vectors, rate/irate/increase, histogram_quantile, aggregation operators (sum/avg/max/min/topk/bottomk), binary operators with vector matching, label manipulation, subqueries
- **Alerting:** Alerting rules (expr, for, labels, annotations), Alertmanager (routing tree, grouping, inhibition, silencing), receivers (Slack, PagerDuty, email, webhook, OpsGenie), Go templating in notifications
- **Configuration:** Scrape configs, relabeling (relabel_configs and metric_relabel_configs), recording rules, OTLP receiver (3.x), global settings, TLS/auth
- **3.x Features:** Native histograms (exponential buckets, server-side quantiles), UTF-8 metric names, OTLP ingestion endpoint, modernized UI

## Prometheus Architecture Quick Reference

```
┌──────────────────────────────────────────────────────────┐
│                    Prometheus Server                      │
│                                                          │
│  ┌──────────┐   ┌──────────┐   ┌──────────────────────┐ │
│  │ Scraper  │   │  TSDB    │   │  HTTP API / PromQL   │ │
│  │ + SD     │──>│  (WAL +  │<──│  + Expression        │ │
│  │          │   │  Blocks) │   │    Browser            │ │
│  └──────────┘   └──────────┘   └──────────────────────┘ │
│       ^                               ^                  │
│       |                               |                  │
│  ┌──────────┐                 ┌──────────────────────┐   │
│  │ Service  │                 │  Rule Manager        │   │
│  │ Discovery│                 │  (Recording + Alert) │   │
│  └──────────┘                 └──────────────────────┘   │
└──────────────────────────────────────────────────────────┘
       ^                               |
       | scrape /metrics               | fire alerts
  ┌────┴─────┐                 ┌───────v────────┐
  │ Targets  │                 │ Alertmanager   │
  └──────────┘                 └────────────────┘
```

**Pull model:** Prometheus scrapes targets at configured intervals. Each target exposes metrics at an HTTP endpoint (default `/metrics`). Targets are discovered dynamically via service discovery or configured statically.

**TSDB:** Custom time-series database. Samples stored in 2-hour blocks. Head block in memory + WAL for crash recovery. XOR compression achieves ~1.3 bytes/sample. Retention by time (default 15d) or size.

## PromQL Quick Reference

### Essential Functions

| Function | Use | Example |
|----------|-----|---------|
| `rate(v[d])` | Per-second rate of counter increase | `rate(http_requests_total[5m])` |
| `increase(v[d])` | Total increase over range | `increase(errors_total[1h])` |
| `histogram_quantile(q, v)` | Percentile from histogram | `histogram_quantile(0.99, sum by (le) (rate(duration_bucket[5m])))` |
| `sum by (l) (v)` | Sum grouped by label | `sum by (job) (rate(requests_total[5m]))` |
| `topk(k, v)` | Top k series by value | `topk(5, sum by (handler) (rate(requests[5m])))` |
| `absent(v)` | Returns 1 if metric missing | `absent(up{job="api"})` |
| `predict_linear(v[d], t)` | Linear prediction | `predict_linear(disk_avail[6h], 4*3600) < 0` |
| `label_replace(v, ...)` | Regex label manipulation | `label_replace(up, "svc", "$1", "instance", "(.+):\\d+")` |

### Common Patterns

**Error rate percentage:**
```promql
100 * sum(rate(http_requests_total{status=~"5.."}[5m])) by (job)
  / sum(rate(http_requests_total[5m])) by (job)
```

**99th percentile latency:**
```promql
histogram_quantile(0.99,
  sum by (le, job) (rate(http_request_duration_seconds_bucket[5m]))
)
```

**CPU utilization per node:**
```promql
1 - avg by (instance) (rate(node_cpu_seconds_total{mode="idle"}[5m]))
```

**Disk fill prediction:**
```promql
predict_linear(node_filesystem_avail_bytes{fstype!~"tmpfs|overlay"}[6h], 4*3600) < 0
```

## Top 10 Operational Rules

1. **Monitor cardinality** -- Check `prometheus_tsdb_head_series` and TSDB status page (`/tsdb-status`) regularly. Each series costs 3-6 KB RAM.

2. **Use recording rules for dashboards** -- Pre-compute expensive `rate()` and `histogram_quantile()` expressions. Name them `level:metric:operations`.

3. **Set `for` on every alert** -- Never fire immediately on transient spikes. Minimum `for: 1m` for critical, `for: 5m` for warnings.

4. **Use `metric_relabel_configs` to drop noise** -- Drop high-cardinality labels and unused metrics at scrape time, not at query time.

5. **Configure remote_write for retention > 30 days** -- Prometheus local storage is not designed for long-term. Use Thanos, Mimir, or VictoriaMetrics.

6. **Run Alertmanager as a cluster (3 nodes)** -- Alertmanager uses gossip for deduplication. Single-node Alertmanager is a SPOF.

7. **Use native histograms (3.x)** -- Exponential buckets eliminate bucket configuration, reduce cardinality, and enable accurate server-side quantiles.

8. **Never put unbounded values in labels** -- User IDs, request IDs, email addresses, and full URLs cause cardinality explosion.

9. **Set `sample_limit` per scrape job** -- Prevent a single target from overwhelming Prometheus with unexpected cardinality.

10. **Use Kubernetes service discovery with relabeling** -- Annotate pods with `prometheus.io/scrape: "true"` and filter in `relabel_configs`.

## Common Pitfalls

**1. Using `irate()` in alerting rules**
`irate()` uses only the last two samples and is too volatile for alerting. Use `rate()` for alerts (smoother, more reliable). Use `irate()` only for dashboards showing momentary spikes.

**2. Forgetting `by (le)` in histogram_quantile**
`histogram_quantile(0.99, rate(http_request_duration_seconds_bucket[5m]))` fails without grouping by `le`. Always include `le` in the aggregation.

**3. Scraping too frequently**
Scraping every 5s generates 3x more data than 15s. Default 15s is appropriate for most use cases. Only reduce for real-time dashboards on specific targets.

**4. Not setting retention limits**
Default retention is 15 days. Without `--storage.tsdb.retention.size`, disk can fill during traffic spikes. Set both time and size retention.

**5. Using summaries instead of histograms**
Summaries compute quantiles client-side and cannot be aggregated across instances. Histograms allow server-side aggregation. Always prefer histograms.

**6. Treating Prometheus as long-term storage**
Prometheus is optimized for recent data. For queries spanning months or years, use remote_write to a dedicated long-term store (Thanos, Mimir, VictoriaMetrics).

**7. Single Prometheus instance in production**
Run two identical Prometheus instances scraping the same targets. Use Alertmanager cluster for alert deduplication. Use Thanos/Mimir Query for query deduplication.

## Storage Sizing Reference

**Formula:**
```
disk_bytes = (active_series / scrape_interval_seconds) * 1.3 bytes * retention_seconds
```

| Active Series | Scrape Interval | 15-day Retention | 30-day Retention |
|--------------|----------------|-----------------|-----------------|
| 100,000 | 15s | ~7 GB | ~15 GB |
| 500,000 | 15s | ~37 GB | ~75 GB |
| 1,000,000 | 15s | ~75 GB | ~150 GB |
| 5,000,000 | 15s | ~375 GB | ~750 GB |

Add 20% overhead for WAL and compaction scratch space. Use SSD for the WAL directory.

## Reference Files

Load these when you need deep knowledge for a specific area:

- `references/architecture.md` -- Pull model design, TSDB internals (WAL, blocks, compaction, retention), service discovery mechanisms, native histograms, UTF-8 names, OTLP ingestion, federation, remote read/write, Thanos/Mimir integration patterns, HA configurations. Read for architecture and scaling questions.
- `references/promql.md` -- Data types (instant/range vectors, scalars), selectors and matchers, rate functions, histogram functions, aggregation operators, binary operators and vector matching, label manipulation, utility functions, 17+ real-world query examples. Read for any PromQL question.
- `references/alerting.md` -- Alert rule structure and examples, Alertmanager architecture, configuration (routing, grouping, inhibition, silencing), receiver types (Slack, PagerDuty, email, webhook), Go templating, alert fatigue prevention. Read for alerting questions.
- `references/configuration.md` -- prometheus.yml structure, global settings, scrape_configs options, Kubernetes SD scrape config, relabeling patterns (relabel_configs and metric_relabel_configs), recording rules, OTLP receiver config, remote write/read. Read for configuration questions.
- `references/diagnostics.md` -- TSDB status analysis, cardinality investigation, scrape failure diagnosis, high memory troubleshooting, high disk usage, slow query optimization, WAL corruption recovery, Alertmanager diagnostics, key internal metrics. Read for troubleshooting.

## Diagnostic Scripts

Ready-made HTTP API scripts (read-only) in `scripts/`, numbered by investigation order.

- `scripts/01-health-and-targets.sh` -- Down targets with errors, TSDB head stats
- `scripts/02-cardinality-audit.sh` -- Top metrics/label-pairs by series count (the explosion finder)
- `scripts/03-rule-health.sh` -- Broken/slow rules and currently firing alerts

Attribution

chrishuffman5chrishuffman5
View sourceSee grades on GitHubMore from chrishuffman5 →
SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments (0)

No comments yet. Be the first to comment!

SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Related Skills

Terraform Module Library

Build reusable Terraform modules for AWS, Azure, and GCP infrastructure following infrastructure-as-code best practices. Use when creating infrastructure modules, standardizing cloud provisioning, or implementing reusable IaC components.

401991 votes

sematext-otel

Wire a service's OpenTelemetry output to Sematext Cloud. Walks through region, App-type, instrumentation flow (managed OTLP endpoint vs Sematext Agent), and signal selection (traces/metrics/logs), then produces the exact env-var block and points at a runnable reference example in this repo. Invoke when instrumenting a new app for Sematext.

01 votes

Deployment Patterns

Deployment workflows, CI/CD pipeline patterns, Docker containerization, health checks, rollback strategies, and production readiness checklists for web applications. Use when setting up deployment infrastructure or planning releases.

2699140 votes

Babysit

Watch a pull request or review cycle until it is ready to merge. Use when asked to babysit, monitor, or keep checking PR comments, reviews, and CI until all actionable issues are resolved.

971540 votes

V7 Roster

Interact with the Paperclip control plane API for task coordination and governance. Use when checking assignments, updating issue status, posting comments, delegating work, managing routines, or calling Paperclip API endpoints.

953190 votes
View all in devops →