Extracts an answer from logs, traces, or metrics — finding the relevant lines in volume, correlating across services, and telling signal from noise. Use this whenever the user points at a log file, asks what happened at a particular time, mentions grepping logs, wants to know how often something occurs, or is trying to reconstruct a sequence of events across services. For fixing what the logs reveal, use debugging; for the write-up afterwards, use root-cause-analysis.
Scanned 9/3/2026
Install to Claude Code
npx -y skills add OKHP3/skillz --skill log-analysis --agent claude-codeInstalls into .claude/skills of the current project.
Are you the author of Log Analysis?
Add the live security badge to your README — it updates automatically with every re-scan.
[](https://www.skillsdirectory.com/skills/okhp3-log-analysis)More formats (shields.io, HTML) on the badges page.
---
name: log-analysis
description: Extracts an answer from logs, traces, or metrics — finding the relevant lines in volume, correlating across services, and telling signal from noise. Use this whenever the user points at a log file, asks what happened at a particular time, mentions grepping logs, wants to know how often something occurs, or is trying to reconstruct a sequence of events across services. For fixing what the logs reveal, use debugging; for the write-up afterwards, use root-cause-analysis.
license: MIT
---
# Log analysis
Logs are a haystack that grows faster than you can read it. The skill is not reading logs — it is
constructing a query narrow enough to answer one question, then widening only as far as needed.
The failure this prevents: scrolling. Scrolling through logs feels like work and finds only what
happens to be near the cursor.
## 1. Ask one answerable question
Before opening anything, write the question down. "What went wrong?" is not answerable. These
are:
- Did request `abc-123` reach the payment service?
- How many 500s between 14:00 and 14:30, and on which endpoint?
- What is the first error after the deploy at 13:47?
- Which tenant accounts for the spike?
**Done when:** you have a question with a checkable answer.
## 2. Anchor on time and identity
Two anchors make everything else tractable:
- **A time window:** bound it tightly, then widen. Start a few minutes before the first known
symptom, because the cause usually precedes it.
- **An identifier:** request ID, trace ID, user, order, tenant. One identifier that threads
through services turns a search into a story.
If there is no correlating ID, that is your most important finding. Nothing else you do here
will be reliable, and adding one should be the follow-up action.
**Done when:** you have a window and, ideally, an ID to follow.
## 3. Cut volume before reading
Filter, then aggregate, then read. Reading first is what wastes the afternoon.
```bash
# Shape of the problem before any individual line
grep ERROR app.log | awk '{print $5}' | sort | uniq -c | sort -rn | head
# Rate over time — is it constant, a spike, or a step change?
grep ERROR app.log | cut -c1-16 | uniq -c
# Follow one request across a file
grep 'req_id=abc-123' *.log | sort -k1,2
```
For structured logs, use the query language rather than grep — `jq` locally, or the platform's
own filtering. Structured logs exist so you can aggregate; grepping them wastes that.
**Done when:** you know the shape — how many, how often, since when, affecting whom.
## 4. Read the boundaries of the incident
The most informative lines are rarely the loudest.
- **The first occurrence.** Not the loudest error, the earliest one. Errors cascade, and the
hundred downstream failures are noise around one upstream cause.
- **The last normal line** before it started, and what immediately follows it.
- **What stopped appearing.** A log line that vanishes is as meaningful as one that appears — a
heartbeat that stopped, a job that never logged completion.
- **The gap.** Silence in a normally chatty service usually means blocked, not idle.
**Done when:** you can state the first symptom and what preceded it.
## 5. Correlate before concluding
- Line up the timeline against deploys, config changes, feature flag flips, scaling events, and
scheduled jobs. Most incidents correlate with a change.
- Compare the affected population to an unaffected one — same time, different region or version.
A natural control is worth more than any amount of reading.
- **Check the clocks.** Servers in different timezones, or logs in local time and UTC mixed, will
produce a false ordering and a wrong conclusion. Verify before trusting sequence.
**Done when:** the sequence is confirmed by more than one source.
## 6. Report what the logs support, and no more
Logs show what was recorded, which is not the same as what happened. Be explicit about the gap:
- **Say what you searched:** the window, the query, the sources. A finding without its query
cannot be checked or repeated.
- **Distinguish absence of evidence from evidence of absence.** "No error logged" may mean it
did not happen, or that the path has no logging, or that logs were dropped under load. Say
which you believe and why.
- **Note sampling and retention.** Sampled traces and rotated logs both hide things.
- **Quote line counts, not impressions.** "412 occurrences across 3 hosts" beats "lots".
## Improving what you found
Every log investigation exposes a gap. Note them as follow-ups: a missing correlation ID, an
error logged without context, a swallowed exception, a log at the wrong level. Fixing those is
what makes the next incident shorter, and it is the most valuable output of this work.
Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.
No comments yet. Be the first to comment!