Skills DirectorySkills Directory
SkillsLearnSecurityCategoriesDocsCommunityBlog
Sign InSubmit Skill
Skills Directory

Security-tested agent skills for Claude, coding agents, and AI workflows.

Directory

  • Browse Skills
  • All Skills A–Z
  • Claude Skills
  • Claude Code Skills
  • Agent Skills
  • Categories
  • Submit a Skill

Learn

  • Learn Hub
  • Install Claude Skills
  • Write SKILL.md
  • Skills vs MCP
  • Directories Compared

Security

  • Security
  • Methodology
  • Secure Claude Skills
  • Security Badges

Company

  • About
  • Community
  • Blog
  • API Docs
  • Advertise

2026 Skills Directory. All rights reserved.

Back to skills

Tail Latency Analysis

ASecurity

Diagnosing and mitigating end-to-end latency tails: defining the latency population, decomposing stage and queue time with per-request evidence, quantifying fan-out under dependence, attributing correlated JVM/OS/network/dependency events, and selecting bounded tail-tolerance mechanisms such as deadlines, partial results, hedging and load-aware routing. Use when p99/p99.9 regresses, stage percentiles do not explain an end-to-end percentile, deploys create cold tails, wide fan-out amplifies ra...

2 stars
0 votes
0 copies
0 views
Added 9/19/2026
developmentgojavanodegitapi

Works with

cliapi

Security Analysis

A100/100

Scanned 9/19/2026

Install to Claude Code

$npx -y skills add robsonkades/agent-skills --skill tail-latency-analysis --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Tail Latency Analysis?

Add the live security badge to your README — it updates automatically with every re-scan.

Security grade badge for Tail Latency Analysis
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/robsonkades-tail-latency-analysis/badge)](https://www.skillsdirectory.com/skills/robsonkades-tail-latency-analysis)

More formats (shields.io, HTML) on the badges page.

Download Zip
Files
SKILL.md
---
name: tail-latency-analysis
description: >
  Diagnosing and mitigating end-to-end latency tails: defining the latency population,
  decomposing stage and queue time with per-request evidence, quantifying fan-out under
  dependence, attributing correlated JVM/OS/network/dependency events, and selecting
  bounded tail-tolerance mechanisms such as deadlines, partial results, hedging and
  load-aware routing. Use when p99/p99.9 regresses, stage percentiles do not explain an
  end-to-end percentile, deploys create cold tails, wide fan-out amplifies rare stragglers,
  or a hedge/retry is proposed. Percentile estimation belongs to latency-statistics;
  queueing models to queueing-models; collector and OS mechanisms to their owning skills.
---

# Tail Latency Analysis

## Purpose

Identify which requests are slow, where their elapsed time went, what condition caused it,
and which intervention improves user outcomes without destabilizing the system.

Tail latency is conditional on endpoint, outcome, tenant, payload, time, load, topology and
deadline. A global p99 can move because the mixture changed even when every cohort stayed
constant (or hide a cohort regression). Start from a precisely scoped population.

## Workflow

Use the steps needed for the requested explanation, diagnosis or policy change. Reuse adequate
evidence and retain a sound design. A narrow answer needs its assumptions, supported conclusion
and relevant limits; it need not produce a full cohort report or new load/failure campaign.
Missing deployment data limits deployment claims, not independent mathematical conclusions.

### 1. Define the user objective

State latency start/end events, population, success/error treatment, deadline/censoring,
window and aggregation. Select quantiles from user impact and available sample precision;
do not mandate p99, p99.9 and maximum for every service. Maximum is highly sample-size and
duration dependent and is useful as an incident exemplar, not a stable SLO statistic.

For runtime-specific claims or changes, inspect deployed Java/JDK, client/telemetry versions
and effective configuration; no upgrade is implied. Missing recordings or request outcomes
remain unknown. Use existing authorization
for bounded captures and isolated or authorized load/failure experiments.

Record offered, admitted and successful work. Closed-loop or completion-only measurements
can underrepresent the worst intervals; check coordinated omission and timed-out/abandoned
requests first. A naturally completion-paced population can legitimately use a closed-loop
workload; state that scope rather than extrapolating it to independently arriving traffic.

### 2. Segment before attributing

Select relevant comparisons by endpoint/operation, payload or work size, tenant/partition,
outcome, instance/node/zone, cache state, deployment age, load and time. Keep cardinality
bounded in metrics; use traces/exemplars or offline joins for high-cardinality dimensions.

Mixture decomposition distinguishes:

- a larger fraction of an existing slow cohort;
- an unchanged fraction whose conditional latency worsened;
- a new slow path;
- a global correlated event;
- estimator/instrumentation change.

Multimodality can suggest distinct paths, but the absence of modes does not prove queueing,
and a component count does not identify causes.

### 3. Decompose per request, not by percentile arithmetic

For one request, if instrumentation partitions elapsed time into disjoint intervals covering
the chosen boundary, an accounting model is:

\[
T_{end}=T_{client}+T_{network}+T_{admission}+T_{queue}
+T_{service}+T_{downstream}+T_{serialization}
\]

Actual spans often overlap or nest, so the displayed sum is not valid on raw span durations.
Reconstruct critical-path spans and waits for sampled slow requests. Per-stage
histograms localize candidates, but neither matching stage p99 nor summing stage p99 proves
ownership: different requests can occupy each percentile and stages can correlate.

Use [decomposing the tail](references/decomposing-the-tail.md).

### 4. Quantify fan-out with dependence explicit

For \(N\) identically distributed parallel leaves, if each independently exceeds threshold
\(t\) with probability \(p\):

\[
P(\max_i T_i>t)=1-(1-p)^N
\]

For 100 leaves and \(p=0.01\), the probability is about 63.4%. Independence is a scenario,
not a default: shared dependencies, synchronized pauses and common requests create positive
dependence; load balancing and mutually exclusive paths can change it differently.

Estimate joint behavior from request-level traces or bounds. A user budget can be allocated
backward only after topology, quorum/partial-result rule and dependence assumptions are
declared. Sequential latency is a sum; parallel latency may be a maximum, order statistic
or deadline-limited partial result. The maximum of leaf durations models all-of-N only when
launch times align; otherwise include launch offsets and parent/merge/cleanup overhead.

### 5. Correlate candidate causes

For an attribution claim, align slow requests with the relevant queue/admission, useful-load,
JVM, process/container, network/storage or dependency evidence. Evidence must overlap the
affected interval and instance. Co-occurrence alone is not causation; compare unaffected
instances/cohorts and perform a controlled change when possible.

Use [attributing the tail](references/attributing-the-tail.md).

### 6. Select a mechanism from cause and topology

Prefer source removal—reduce contention, queueing, pauses, skew or expensive work—when
feasible. Tail-tolerance mechanisms trade extra work, completeness, errors, state and
complexity:

| Cause/topology                                   | Candidate                              | Key risk                             |
| ------------------------------------------------ | -------------------------------------- | ------------------------------------ |
| local transient straggler, spare diverse replica | delayed hedge                          | incident-time load amplification     |
| wide fan-out permits partial answer              | quorum/k-of-N by deadline              | incomplete/biased results            |
| persistent slow replica                          | bounded outlier ejection/probation     | correlated ejection removes capacity |
| head-of-line blocking                            | classes, fair scheduling, work slicing | starvation/complexity                |
| hot key/partition                                | repartition or selective replication   | consistency/rebalance cost           |
| saturated shared dependency                      | admission, shedding, capacity repair   | hedge/retry worsens it               |
| cold rollout instance                            | warm-capacity routing/slow start       | rollout duration/cost                |

See [hedging and tail tolerance](references/hedging-and-tail-tolerance.md).

### 7. Validate the claim and affected system

For a material intervention, compare the original population and workload with the changed
policy, including user tail, success/completeness and affected capacity/fairness limits.
Choose normal, degraded and recovery scenarios that discriminate its risks; reuse adequate
existing evidence. A hedge rollout needs aggregate attempt and residual-work bounds under
broad slowdown, not just a nominal 1% hedge rate. A scoped math or source/API review can close
with its supported result and explicit deployment limits.

## Diagnostic rules

- Component percentiles cannot generally be added or subtracted to recover a sum's percentile.
  With the same population and quantile definition, nonnegative sequential durations give
  `T >= T_i` pointwise, hence `q_p(T) >= q_p(T_i)`. This lower bound does not order `q_p(T)`
  against the **sum** of component quantiles or identify the requests causing the tail.
- A stage percentile equal to end-to-end p99 does not prove the same requests drove both.
- A faster p50 with worse p99 may be a regression, but decision weights come from the SLO
  and user impact—not a universal preference for tails.
- Never silently trim slow observations. Exclude only proven measurement corruption with a
  recorded rule and sensitivity analysis.
- Post-deploy slowness can be JIT/cache/classloading/TLS/connection/data warmup, placement,
  dependency or rollout routing. “JIT disabled” and “always JIT” are both unsupported
  without evidence.
- GC logs show collector activity; safepoint, scheduling and allocation stalls require
  their own evidence. Verify JFR event names/configuration against the exact JDK recording.
- Fixed duration bands are triage hints only. Retransmission timers, cgroup periods,
  storage and GC behavior are configurable and layered.

## Failure modes

| Symptom                                        | Discriminator                                    | Next step                                       |
| ---------------------------------------------- | ------------------------------------------------ | ----------------------------------------------- |
| every in-process stage shifts together         | aligned pause/scheduling/client boundary         | correlate JFR, safepoint and OS timeline        |
| one dependency span dominates only slow traces | dependency cohort/outcome/instance               | inspect its queue, retries and topology         |
| tail grows with fan-out width                  | leaf exceedance correlation and completion rule  | reduce width, partial results or safe tolerance |
| tail begins after rollout                      | age since readiness, compilation/cache/placement | measure warm-capacity ramp                      |
| spikes at high load                            | admitted load, queue age, throttling, pools      | queue/resource diagnosis before hedging         |
| periodic spikes                                | aligned GC/jobs/rotation/network/control cycles  | identify phase and test causal disable/shift    |
| timeout boundary pile-up                       | censored durations and remaining work            | propagate deadlines/cancellation                |

## Anti-patterns

**One percentile per service copied end to end:** ignores topology, population and
dependence. Allocate a user objective through the actual critical path and completion rule.

**Stage-percentile accounting:** hides request identity and covariance. Use per-request
critical paths or joint distributions.

**Hedge at historical p95 and call it 5% overhead:** when the distribution shifts, almost
all calls can cross the fixed delay. Enforce a rolling budget/pushback and test degradation.

**Independent retry/hedge policies at multiple layers:** attempt limits multiply and deadlines
or cancellation can be lost. Prefer one policy owner; coordinated layers can be valid when
they enforce one total-attempt/resource budget and end-to-end deadline. Account for actual
bottom-layer work, including residual and transparent attempts.

**Correlation by dashboard eyeballing:** different clocks/windows and mixture changes
produce false matches. Align raw events and compare controls.

## Cross-skill routing

- latency-statistics: estimator, histogram and sample uncertainty.
- distributed-tracing-design: span boundaries, sampling and exemplars.
- queueing-models / coordinated-omission: waiting and missing arrivals.
- pause-attribution / safepoints / gc-log-analysis: JVM pause mechanism.
- linux-for-jvm / ebpf-for-jvm / tcp-tuning: scheduling and network evidence.
- timeouts-and-deadlines / retries-and-backoff / scatter-gather: policy ownership.

## Authoritative references

- [Dean and Barroso: The Tail at Scale](https://research.google/pubs/the-tail-at-scale/)
- [gRPC: Request hedging](https://grpc.io/docs/guides/request-hedging/)
- [gRPC: Deadlines](https://grpc.io/docs/guides/deadlines/)
- [Google SRE: Addressing cascading failures](https://sre.google/sre-book/addressing-cascading-failures/)
- [OpenJDK 25 HotSpot JFR event definitions](https://github.com/openjdk/jdk/blob/jdk-25%2B36/src/hotspot/share/jfr/metadata/metadata.xml) — native event fields, not recording enablement.
- [OpenJDK 25 default JFR settings](https://github.com/openjdk/jdk/blob/jdk-25%2B36/src/jdk.jfr/share/conf/jfr/default.jfc) — template settings, not proof of a target recording's effective configuration.

Attribution

robsonkadesrobsonkades
View sourceMore from robsonkades →
SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments (0)

No comments yet. Be the first to comment!

SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Related Skills

Browser Extension Developer

Use this skill when developing or maintaining browser extension code in the `browser/` directory, including Chrome/Firefox/Edge compatibility, content scripts, background scripts, or i18n updates.

281612 votes

Seo Optimizer

SEO optimization with keyword analysis, readability assessment, technical validation, content quality. Use for search rankings, blog posts, content audits, or encountering keyword density, readability scores, meta tags, schema markup errors.

2132 votes

Google Official Seo Guide

Official Google SEO guide covering search optimization, best practices, Search Console, crawling, indexing, and improving website search visibility based on official Google documentation

1862 votes

Tanstack Start

Build a full-stack TanStack Start app on Cloudflare Workers from scratch — SSR, file-based routing, server functions, D1+Drizzle, better-auth, Tailwind v4+shadcn/ui. Use whenever the user mentions TanStack Start, asks to scaffold a full-stack Cloudflare app with SSR, wants an SSR dashboard, or asks for a React 19 + Cloudflare Workers app with file-based routing and server functions — even if they don't name TanStack Start specifically. No template repo — Claude generates every file fresh per ...

9881 votes

Pentest

PTES-aligned adversarial security audit for backend, frontend, and mobile applications. Produces a CVSS-scored Hacker Report with verified PoCs and phased remediation.

5491 votes
View all in development →