Pull the most recent failed Langfuse trace (or a specific
Scanned 9/19/2026
Install to Claude Code
npx -y skills add thecoderpanda/fde-starter-kit --skill trace-debug --agent claude-codeInstalls into .claude/skills of the current project.
Are you the author of Trace Debug?
Add the live security badge to your README — it updates automatically with every re-scan.
[](https://www.skillsdirectory.com/skills/thecoderpanda-trace-debug)More formats (shields.io, HTML) on the badges page.
---
name: trace-debug
description: Pull the most recent failed Langfuse trace (or a specific
trace ID), walk the tool calls and completions, and hypothesize a root
cause. Use when the user reports an agent behaved wrong in production
or staging, or when Langfuse shows an error trace.
---
# trace-debug
Purpose: turn a red trace in Langfuse into a specific, actionable
hypothesis in one pass — without opening the browser five times.
## When to fire
- The user says: "why did the agent do X", "debug this trace", "the
agent broke in prod", "trace <uuid>", "look at the failed trace".
- The user pasted a Langfuse URL or trace ID.
## Preconditions
1. `LANGFUSE_PUBLIC_KEY`, `LANGFUSE_SECRET_KEY`, and `LANGFUSE_BASE_URL`
must be set. If any is missing, stop and tell the user to run
`npm run langfuse:up` and populate `.env`.
2. `curl` and `jq` are available. If not, fall back to `node -e` using
`fetch`.
## Steps
1. **Resolve which trace.** If the user gave an ID or URL, extract it.
Otherwise, list the 5 most recent traces where any observation had
`level = "ERROR"`:
```bash
curl -s -u "$LANGFUSE_PUBLIC_KEY:$LANGFUSE_SECRET_KEY" \
"$LANGFUSE_BASE_URL/api/public/traces?limit=5&orderBy=timestamp.desc" \
| jq '.data[] | {id, name, timestamp, metadata}'
```
Ask the user which one — or, if only one, proceed.
2. **Fetch the trace tree.**
```bash
curl -s -u "$LANGFUSE_PUBLIC_KEY:$LANGFUSE_SECRET_KEY" \
"$LANGFUSE_BASE_URL/api/public/traces/$TRACE_ID" | jq .
```
You want the ordered `observations[]`: each is an LLM call or a tool
call with input, output, and any error string.
3. **Walk the tree top to bottom.** For each observation, produce one
line:
```
<t+ms> <type> <name> <status> <one-line input> <one-line output/error>
```
Truncate strings to ~120 chars. This is the timeline the user reads
first.
4. **Identify the failure point.** The first observation with a non-null
`error` or a tool result containing `{ "error": ... }` is usually
ground zero. If none, the failure is a bad LLM completion — look at
the last generation's `output`.
5. **Correlate to code.** Map the failing observation to the source:
- Tool name → `./agent/tools/<name>.ts`. Read the `execute` body.
- LLM completion → check the prompt (`./agent/prompts/system.ts`)
and the tool schema the model was calling.
6. **Hypothesize.** State one specific hypothesis in the form
*"the failure happened because X, evidence: Y, fix: Z"*. If the
evidence is thin, say so and list what you'd need to be certain.
## Definition of done
- Timeline printed with one line per observation.
- Failure point identified with a file:line reference where possible.
- One hypothesis with evidence and a proposed next step.
- You did NOT apply the fix — this skill only diagnoses.
Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.
No comments yet. Be the first to comment!