Skills DirectorySkills Directory
SkillsLearnSecurityCategoriesDocsBlogPro
Sign InSubmit Skill
Skills Directory

Security-tested agent skills for Claude, coding agents, and AI workflows.

Directory

  • Browse Skills
  • All Skills A–Z
  • Claude Skills
  • Claude Code Skills
  • Agent Skills
  • Categories
  • Authors
  • Submit a Skill

Learn

  • Learn Hub
  • Install Claude Skills
  • Write SKILL.md
  • Skills vs MCP
  • Directories Compared

Security

  • Security
  • Methodology
  • Secure Claude Skills
  • Security Badges
  • Chrome Extension
  • Skill Manager

Company

  • About
  • Community
  • Blog
  • API Docs
  • Advertise

2026 Skills Directory. All rights reserved.

ProTermsPrivacyRefunds
Back to skills

Sre Slos

ASecurity

SRE SLI/SLO/SLA implementation

2 stars
0 votes
0 copies
0 views
Added 9/29/2026
ai-agentspythongoapidevops

Works with

cliapi

Security Analysis

A100/100

Pro scans all 2 files and shows the line behind each finding

Scanned 9/29/2026

$npx -y skills add ssrjkk/claude-skills --skill sre-slos --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Sre Slos?

Add the live security badge to your README — it updates automatically with every re-scan.

Security grade badge for Sre Slos
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/ssrjkk-sre-slos/badge)](https://www.skillsdirectory.com/skills/ssrjkk-sre-slos)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
Files
SKILL.md
---
name: sre-slos
description: "SRE SLI/SLO/SLA implementation"
category: devops
tags: [sre, sli, slo, sla, reliability, monitoring]
models: [sonnet, opus]
version: 1.0.0
created: 2026-05-14
updated: 2026-09-06
---
# SRE SLOs

> Implement Service Level Indicators, Objectives, and Agreements following SRE best practices.

## Quick Start
```yaml
# slo-config.yaml — Service level configuration
apiVersion: sre.google.com/v1
kind: SLO
metadata:
  name: api-availability
  service: payment-api
spec:
  description: "Payment API availability SLO"
  target: 99.9  # percentage
  window: 28d   # rolling window
  
  indicator:
    type: availability
    definition: |
      # SLI: Ratio of successful requests
      good_events = count(status_code < 500)
      valid_events = count(status_code != 0)
      sli = good_events / valid_events
  
  burnRateAlerts:
    - severity: page
      threshold: 0.01  # minutes of error budget consumed per minute
      lookback: 1h
    - severity: ticket
      threshold: 0.001
      lookback: 6h
---
# Error budget policy
apiVersion: sre.google.com/v1
kind: ErrorBudget
metadata:
  name: api-error-budget
spec:
  sloRef: api-availability
  policy:
    # Stop deployments when error budget is depleted
    deployFreeze:
      enabled: true
      remainingBudgetPercent: 20
```

```python
# SLO Monitoring with Prometheus
from prometheus_client import Histogram, Counter
import time

REQUEST_LATENCY = Histogram(
    'http_request_duration_seconds',
    'HTTP request latency',
    ['method', 'endpoint'],
    buckets=[0.01, 0.05, 0.1, 0.25, 0.5, 1.0, 2.5, 5.0]
)

REQUEST_COUNT = Counter(
    'http_requests_total',
    'Total HTTP requests',
    ['method', 'endpoint', 'status']
)

def track_request(method: str, endpoint: str):
    def decorator(func):
        def wrapper(*args, **kwargs):
            start = time.time()
            try:
                result = func(*args, **kwargs)
                REQUEST_COUNT.labels(method=method, endpoint=endpoint, status="200").inc()
                return result
            except Exception:
                REQUEST_COUNT.labels(method=method, endpoint=endpoint, status="500").inc()
                raise
            finally:
                REQUEST_LATENCY.labels(method=method, endpoint=endpoint).observe(time.time() - start)
        return wrapper
    return decorator
```

## Key Concepts
SLIs measure service reliability (latency, availability, durability). SLOs set targets (e.g., 99.9% availability over 28 days). Error budget = (1 - SLO) × total events. Use burn rate alerts for early detection.

## When to Use
- Defining reliability expectations for services
- Making data-driven decisions about deployment velocity
- Balancing feature development with reliability investment

## Step-by-Step
1. Pick service boundaries: define the user-visible promise (e.g., "API replies within 300ms p50, 99.9% available").
2. Define SLIs as ratios: good events / valid events, per endpoint, aggregated over a rolling window.
3. Set SLO targets with error budgets: availability = 99.9% over 28 days → budget = 43 min of downtime.
4. Wire monitoring: export latency histograms and counters from Prometheus; compute SLIs with recording rules.
5. Add burn-rate alerts: page on 2h lookback at 14.4x budget consumption, ticket on 6h/1d windows.
6. Gate change velocity: pause deployments when remaining budget crosses the policy threshold (e.g., 20%).

## Examples
```yaml
# Multi-window burn-rate alert (two tiers)
groups:
  - name: slo-alerts
    rules:
      - alert: APIAvailabilityBurnRate
        expr: |
          sum(rate(http_requests_total{status=~"5.."}[2h]))
          / sum(rate(http_requests_total[2h])) > 0.02
        # 2h window, ~14.4x budget consumption at 99.9% SLO
        annotations:
          summary: "API availability error budget burning fast"
```
```python
# Compute SLO compliance from a Prometheus Query API response
from prometheus_api_client import PrometheusConnect
p = PrometheusConnect(url="http://prometheus:9090")
sli_good = p.custom_query('sum(rate(http_requests_total{status<"500"}[28d]))')
sli_valid = p.custom_query('sum(rate(http_requests_total[28d]))')
print("availability:", float(sli_good[0]["value"][1]) / float(sli_valid[0]["value"][1]))
```

## Validation
1. SLIs are accurately measured and reported
2. SLO compliance dashboard shows current and historical status
3. Error budget alerts fire correctly during degradation
4. Deployment gates respect error budget policy

Attribution

ssrjkkssrjkk
View sourceSee grades on GitHubMore from ssrjkk →
SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments (0)

No comments yet. Be the first to comment!

SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Related Skills

Caveman

Terse caveman voice: answer first, fluff gone, every technical fact kept. Use for /caveman, "caveman mode", "talk like caveman", "be brief", "less tokens". Stays on until "stop caveman" or "normal mode".

1100021 votes

Hyperplan

Adversarial multi-agent planning skill. Self-orchestrates 5 hostile category members (unspecified-low, unspecified-high, deep, ultrabrain, artistry) via team-mode for ruthless cross-critique debate, distills only the defensible insights, then MANDATORILY hands the distilled insight bundle to the `plan` agent for executable plan formalization. Use when planning needs maximum rigor and surfacing of weak assumptions, blind spots, and over-engineering. Triggers: 'hyperplan', 'hpp', '/hyperplan', ...

698461 votes

Writing Skills

Create and manage Claude Code skills in HASH repository following Anthropic best practices. Use when creating new skills, modifying skill-rules.json, understanding trigger patterns, working with hooks, debugging skill activation, or implementing progressive disclosure. Covers skill structure, YAML frontmatter, trigger types (keywords, intent patterns), UserPromptSubmit hook, and the 500-line rule. Includes validation and debugging with SKILL_DEBUG. Examples include rust-error-stack, cargo-dep...

3931 votes

Mcp Code Execution

Routes multi-tool workflows through MCP servers for large datasets and pipelines. Use when Bash tool overhead is limiting throughput on data-heavy tasks.

3421 votes

catchup

Recovers the conversation and failed tool calls of a previous Codex, Amp, Claude Code, Antigravity, Cline, Copilot CLI, Cursor, DeepSeek Harness, Grok Build, Kimi, OpenCode, Pi Agent, or ZCode session. Use when the user says "catch up", "what did the last session do", "get me up to speed", "I switched agents", asks to recover/summarize a previous session before continuing, or asks to diagnose or report a catchup failure. Do NOT use for the current conversation, git history, or any non-agent log.

741 votes
View all in ai-agents →