Skills DirectorySkills Directory
SkillsLearnSecurityCategoriesDocsBlogPro
Sign InSubmit Skill
Skills Directory

Security-tested agent skills for Claude, coding agents, and AI workflows.

Directory

  • Browse Skills
  • All Skills A–Z
  • Claude Skills
  • Claude Code Skills
  • Agent Skills
  • Categories
  • Authors
  • Submit a Skill

Learn

  • Learn Hub
  • Install Claude Skills
  • Write SKILL.md
  • Skills vs MCP
  • Directories Compared

Security

  • Security
  • Methodology
  • Secure Claude Skills
  • Security Badges
  • Chrome Extension
  • Skill Manager

Company

  • About
  • Community
  • Blog
  • API Docs
  • Advertise

2026 Skills Directory. All rights reserved.

ProTermsPrivacyRefunds
Back to skills

Test Result Analyzer

ASecurity

Use when Ingests test logs and identifies root causes across multiple failing test files. Provides actionable fix recommendations.

5 stars
0 votes
0 copies
0 views
Added 9/27/2026
ai-agentsgobashnodeexpresstestingdebuggingapidatabaseci/cd

Works with

terminalapi

Security Analysis

A96/100
mediumInstalls packages at runtime which could introduce malicious dependencies

Pro shows the line behind each finding and how to fix it

Scanned 9/27/2026

$npx -y skills add Harmitx7/tribunal-kit --skill test-result-analyzer --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Test Result Analyzer?

Add the live security badge to your README — it updates automatically with every re-scan.

Security grade badge for Test Result Analyzer
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/harmitx7-test-result-analyzer-tribunal-kit/badge)](https://www.skillsdirectory.com/skills/harmitx7-test-result-analyzer-tribunal-kit)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
Files
SKILL.md
---
name: test-result-analyzer
description: "Use when Ingests test logs and identifies root causes across multiple failing test files. Provides actionable fix recommendations."
version: 5.0.0
last-updated: 2026-09-13
skills:
  - systematic-debugging
  - testing-patterns
  - tdd-workflow
tools: Read, Grep, Glob, Bash, Edit, Write
scripts-binding:
  - .agent/scripts/test_runner.js
  - .agent/scripts/verify_all.js
  - .agent/scripts/lint_runner.js
---

# Test Result Analyzer Skill

---

## šŸ› ļø Technical Architecture & Reference Recipes

---

## When to Activate

- After a test run with multiple failures.
- When the user says "tests are failing", "analyze test results", "what broke?", or "test failed".
- During CI/CD pipeline debugging.
- When `test_runner.js` or any test command exits with failures.
- When paired with `systematic-debugging` for deep root-cause investigation.

## Analysis Pipeline

```
Test output (terminal or log file)
    │
    ā–¼
Runner detection — identify test framework from output format
    │
    ā–¼
Failure extraction — parse each FAIL block into structured data
    │
    ā–¼
Clustering — group failures by root module, error type, shared dependency
    │
    ā–¼
FPF detection — find the First Point of Failure
    │
    ā–¼
Dependency graph — map cascade relationships
    │
    ā–¼
Fix recommendations — ordered by impact (most failures resolved first)
    │
    ā–¼
Report — structured output with confidence levels
```

## Step 1: Runner Detection

Auto-detect the test framework from output patterns:

| Framework   | Detection Pattern                                 | Failure Marker                |
| ----------- | ------------------------------------------------- | ----------------------------- |
| Jest        | `PASS`/`FAIL` with file paths, `ā—` for test names | `FAIL src/...`                |
| Vitest      | `āœ“`/`Ɨ` markers, `FAIL` blocks                    | `āÆ FAIL` or `Ɨ test name`     |
| pytest      | `PASSED`/`FAILED` with `::` separator             | `FAILED tests/...::test_name` |
| Go test     | `ok`/`FAIL` with package paths                    | `--- FAIL: TestName`          |
| Mocha       | `passing`/`failing` counts, indented suites       | `N failing` section           |
| JUnit (XML) | `<testsuite>` XML structure                       | `<failure>` elements          |
| RSpec       | `.F` markers, `Failures:` section                 | `Failure/Error:`              |
| Cargo test  | `test result: FAILED`                             | `---- test_name stdout ----`  |

## Step 2: Failure Extraction

For each failure, extract a structured record:

```
{
  test_name:    "should return 401 for unauthenticated requests"
  test_file:    "src/api/auth.test.ts"
  test_line:    42
  error_type:   "AssertionError"
  expected:     "401"
  received:     "200"
  stack_trace:  ["auth.test.ts:42", "auth.middleware.ts:18", "express/router.ts:..."]
  source_files: ["auth.middleware.ts:18"]  // files from YOUR codebase in the stack
}
```

## Step 3: Failure Clustering

Group failures into clusters based on shared characteristics:

### Cluster Types

| Cluster Type        | How to Detect                                         | Typical Root Cause                       |
| ------------------- | ----------------------------------------------------- | ---------------------------------------- |
| **Shared Module**   | Multiple tests import from the same file that changed | Missing export, type change, API change  |
| **Same Error Type** | All failures throw `TypeError` or `ConnectionError`   | Broken dependency, env issue             |
| **Shared Fixture**  | Tests using same `beforeEach`/setup fail together     | Fixture setup failure cascading          |
| **Import Chain**    | Failures follow the import graph                      | Dependency that fails to resolve         |
| **Environment**     | All tests fail with connection/config errors          | Missing env var, DB not running          |
| **Timing**          | Tests pass individually, fail together                | Race condition, shared state             |
| **Snapshot**        | Multiple `toMatchSnapshot` failures                   | Intentional UI change (update snapshots) |

### Cascade Detection Algorithm

```
1. Sort failures by file path and execution order.
2. Find the FIRST failure in execution order → candidate FPF.
3. Check if the FPF's source file appears in other failures' import chains.
4. If yes → FPF is the root cause, other failures are cascades.
5. If no → failures are independent (multiple root causes).
```

## Step 4: First Point of Failure (FPF) Detection

The FPF is the most valuable finding — fix it first, and cascading failures resolve automatically.

```
Example:
  12 test files fail.
  11 of them import from `utils/auth.ts`.
  The first failure is in `utils/auth.test.ts` at line 42.
  Error: `generateToken is not exported from './auth'`

  FPF: utils/auth.ts:42 — missing export
  Cascade: 11 other test files fail because they can't import generateToken
  Fix: Add `export { generateToken }` to utils/auth.ts
  Expected resolution: 12 of 12 failures (100%)
```

**FPF Confidence Levels:**

| Confidence | Criteria                                                 |
| ---------- | -------------------------------------------------------- |
| **HIGH**   | Same source file in >50% of failure stack traces         |
| **MEDIUM** | Same error type across multiple test files               |
| **LOW**    | Failures appear independent, multiple root causes likely |

## Step 5: Fix Recommendations

For each cluster, provide actionable fixes:

| Fix Type              | Example                                         | How to Verify                         |
| --------------------- | ----------------------------------------------- | ------------------------------------- |
| **Missing Export**    | `export { fn }` added to module                 | Re-run failing tests                  |
| **Type Mismatch**     | Function signature changed, callers need update | Check callers with `grep_search`      |
| **Stale Mock**        | Mock doesn't match new interface                | Compare mock to actual implementation |
| **Env Variable**      | `.env.test` missing `DATABASE_URL`              | Check `.env.example` vs `.env.test`   |
| **Snapshot Update**   | Intentional UI change                           | Run with `--updateSnapshot` flag      |
| **Race Condition**    | Tests share global state                        | Add isolation or `beforeEach` reset   |
| **Dependency Update** | Package API changed after upgrade               | Check changelog of updated package    |

### Fix Priority Formula

```
Priority = (Tests_Resolved Ɨ 10) + (Confidence_Score Ɨ 5) - (Estimated_Fix_Time_Minutes)

Fix in this order:
1. Highest priority score first
2. If tied, prefer HIGH confidence
3. If still tied, prefer fewer files to change
```

## Report Format

```
━━━ Test Result Analysis ━━━━━━━━━━━━━━━━

Runner:    [Jest / Vitest / pytest / Go / auto-detected]
Total:     48 tests across 12 files
Result:    36 passed | 12 failed | 0 skipped
Duration:  4.2s
Coverage:  78% statements (if available)

━━━ First Point of Failure ━━━━━━━━━━━━━━

šŸ“ utils/auth.test.ts → line 42
   Error:  `generateToken` is not exported from `./auth`
   Type:   ImportError
   Impact: Cascades to 11 other test files

   This is the root cause. Fix this first.

━━━ Failure Clusters ━━━━━━━━━━━━━━━━━━━━

Cluster 1: Missing Export  (11 tests, HIGH confidence)
  Root:       utils/auth.ts:42
  Cascade:    auth.test.ts, users.test.ts, sessions.test.ts, ...
  Fix:        Add `export { generateToken }` to auth.ts
  Resolution: 11 of 12 failures (92%)
  Priority:   ā˜…ā˜…ā˜…ā˜…ā˜… (115 pts)

Cluster 2: Stale Mock  (1 test, MEDIUM confidence)
  Root:       api/users.test.ts:98
  Error:      Expected { name, email, role } but received { name, email }
  Fix:        Add `role: "user"` to mock at line 15
  Resolution: 1 of 12 failures (8%)
  Priority:   ā˜…ā˜…ā˜†ā˜†ā˜† (20 pts)

━━━ Fix Plan ━━━━━━━━━━━━━━━━━━━━━━━━━━━

Step 1: Fix utils/auth.ts export
        → Expected: 11 failures resolved
        → Time: ~2 minutes
        → Run: npx jest utils/auth.test.ts (verify FPF fix)

Step 2: Update mock in api/users.test.ts:15
        → Expected: 1 failure resolved
        → Time: ~1 minute

Step 3: Re-run full suite
        → Expected: all 12 failures resolved (0 remaining)

━━━ Warnings ━━━━━━━━━━━━━━━━━━━━━━━━━━━

āš ļø No test coverage report detected. Consider adding --coverage flag.
āš ļø 3 test files have no assertions (test names end in `.todo`).
```

## Edge Cases

### All Tests Fail

```
If 100% of tests fail → likely environment issue, not code:
  1. Check if dev server / database is running
  2. Check .env.test for missing variables
  3. Check node_modules exists (run npm install)
  4. Check for breaking dependency upgrade in recent commits
```

### Flaky Tests

```
If same test passes on retry → flaky:
  1. Check for shared mutable state between tests
  2. Check for time-dependent assertions
  3. Check for unresolved promises / async leaks
  4. Check for network-dependent tests without mocks
```

### Only Snapshot Tests Fail

```
If only snapshot tests fail → likely intentional UI change:
  1. Review snapshot diffs
  2. If changes are expected: run with --updateSnapshot
  3. If changes are unexpected: check for unintended CSS/component changes
```

## Cross-Skill Integration

| Paired Skill           | Integration Point                                        |
| ---------------------- | -------------------------------------------------------- |
| `systematic-debugging` | Escalate when FPF is unclear → 4-phase debug methodology |
| `testing-patterns`     | Reference when recommending test structure improvements  |
| `workflow-optimizer`   | Flag inefficient test-debug-retest loops                 |

## Anti-Hallucination Guard

- **Only analyze test output that was actually produced** — never generate fake test results.
- **Never invent file paths or line numbers** — only reference what appears in the stack trace.
- **Verify source files exist** before suggesting fixes — use `view_file` or `find_by_name`.
- **Mark uncertainty**: `// UNCERTAIN: log format not fully recognized, manual review recommended`.
- **Never guess at assertion values** — quote exactly what "Expected" and "Received" say in the output.
- **Don't assume test runner** — auto-detect from output format, don't assume Jest.

Attribution

Harmitx7Harmitx7
View sourceSee grades on GitHubMore from Harmitx7 →
SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments (0)

No comments yet. Be the first to comment!

SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Related Skills

Caveman

Terse caveman voice: answer first, fluff gone, every technical fact kept. Use for /caveman, "caveman mode", "talk like caveman", "be brief", "less tokens". Stays on until "stop caveman" or "normal mode".

1100021 votes

Hyperplan

Adversarial multi-agent planning skill. Self-orchestrates 5 hostile category members (unspecified-low, unspecified-high, deep, ultrabrain, artistry) via team-mode for ruthless cross-critique debate, distills only the defensible insights, then MANDATORILY hands the distilled insight bundle to the `plan` agent for executable plan formalization. Use when planning needs maximum rigor and surfacing of weak assumptions, blind spots, and over-engineering. Triggers: 'hyperplan', 'hpp', '/hyperplan', ...

698621 votes

Writing Skills

Create and manage Claude Code skills in HASH repository following Anthropic best practices. Use when creating new skills, modifying skill-rules.json, understanding trigger patterns, working with hooks, debugging skill activation, or implementing progressive disclosure. Covers skill structure, YAML frontmatter, trigger types (keywords, intent patterns), UserPromptSubmit hook, and the 500-line rule. Includes validation and debugging with SKILL_DEBUG. Examples include rust-error-stack, cargo-dep...

3931 votes

Mcp Code Execution

Routes multi-tool workflows through MCP servers for large datasets and pipelines. Use when Bash tool overhead is limiting throughput on data-heavy tasks.

3421 votes

catchup

Recovers the conversation and failed tool calls of a previous Codex, Amp, Claude Code, Antigravity, Cline, Copilot CLI, Cursor, DeepSeek Harness, Grok Build, Kimi, OpenCode, Pi Agent, or ZCode session. Use when the user says "catch up", "what did the last session do", "get me up to speed", "I switched agents", asks to recover/summarize a previous session before continuing, or asks to diagnose or report a catchup failure. Do NOT use for the current conversation, git history, or any non-agent log.

741 votes
View all in ai-agents →