Skills DirectorySkills Directory
SkillsLearnSecurityCategoriesDocsBlogPro
Sign InSubmit Skill
Skills Directory

Security-tested agent skills for Claude, coding agents, and AI workflows.

Directory

  • Browse Skills
  • All Skills A–Z
  • Claude Skills
  • Claude Code Skills
  • Agent Skills
  • Categories
  • Authors
  • Submit a Skill

Learn

  • Learn Hub
  • Install Claude Skills
  • Write SKILL.md
  • Skills vs MCP
  • Directories Compared

Security

  • Security
  • Methodology
  • Secure Claude Skills
  • Security Badges
  • Chrome Extension
  • Skill Manager

Company

  • About
  • Community
  • Blog
  • API Docs
  • Advertise

2026 Skills Directory. All rights reserved.

ProTermsPrivacyRefunds
Back to skills

Mutation Testing

ASecurity

Configures mewt or muton campaigns, analyzes surviving mutants, and investigates bugs exposed by testing gaps. Use when setting up mutation testing, reviewing campaign results, identifying equivalent mutants, or finding bugs from surviving mutations.

7,287 stars
0 votes
0 copies
9 views
Added 9/29/2026
researchjavascriptrustgojavashellbashsqltestingapidatabase

Works with

api

Security Analysis

A100/100

Pro scans all 12 files and shows the line behind each finding

Scanned 9/29/2026

$npx -y skills add trailofbits/skills --skill mutation-testing --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Mutation Testing?

Add the live security badge to your README — it updates automatically with every re-scan.

Security grade badge for Mutation Testing
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/trailofbits-mutation-testing/badge)](https://www.skillsdirectory.com/skills/trailofbits-mutation-testing)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
Files
SKILL.md
---
name: mutation-testing
description: "Configures mewt or muton campaigns, analyzes surviving mutants, and investigates bugs exposed by testing gaps. Use when setting up mutation testing, reviewing campaign results, identifying equivalent mutants, or finding bugs from surviving mutations."
allowed-tools: Read Write Bash Grep

---

# Mutation Testing (mewt/muton)

Routes to the right mutation testing workflow and loads the references that workflow needs.

> **Note**: muton and mewt share identical interfaces. Examples use `mewt`; substitute `muton` and its file names (`muton.toml`, `muton.sqlite`) for muton projects.

`mewt --help` and `mewt <subcommand> --help` are the source of truth for command-line behavior. Examples below reflect the mewt 4.x API; run `--help` when a flag looks unfamiliar or a command fails.

## When to Use

Use this skill when the user:
- Mentions "mewt", "muton", or "mutation testing"
- Wants to configure, scope, or speed up a mutation testing campaign
- Wants to analyze mutation results — surviving/uncaught mutants, equivalent mutants, kill rate
- Wants to use mutation results to find bugs in the source code

## When NOT to Use

Do not use this skill when the user asks about tests or line coverage without any mutation testing context.

---

## Routing

Pick the workflow, then load it together with the references listed for it. Workflows and references do not load each other — that decision belongs here.

**Setting up, scoping, or speeding up a campaign**
→ [workflows/configuration.md](workflows/configuration.md)
→ Also load [references/optimization-strategies.md](references/optimization-strategies.md) when the campaign estimate is long enough to need trimming, or the user asks to make it faster.

**Campaign finished, hunting for bugs in untested code**
→ [workflows/bug-hunter.md](workflows/bug-hunter.md)

**Turning results into a formal analysis report**
→ [workflows/analyzing-results.md](workflows/analyzing-results.md), plus:
- [references/equivalent-mutants.md](references/equivalent-mutants.md) — equivalence catalog and verification procedure
- [references/severity-classification.md](references/severity-classification.md) — severity tier criteria
- [references/report-template.md](references/report-template.md) — report structure
- [references/blockchain-patterns.md](references/blockchain-patterns.md) — **only** for Solidity, Move, FunC/Tolk, Cairo, or Solana Rust targets
- [references/input-formats.md](references/input-formats.md) — unless the results came from mewt or muton. Foreign tool output may not be self-describing; this covers the parsing anchors for slither-mutate, mull, and dextool-mutate

**Anything else** → run `mewt --help` or `mewt <subcommand> --help`, then assist directly.

---

## Essential Commands

```bash
# Set up and run
mewt init                    # Create config and database
mewt mutate [paths]          # Generate mutants without testing them
mewt run [paths]             # Generate mutants and run the campaign

# Read results
mewt status                  # Overview with per-file breakdown
mewt results                 # Uncaught mutants (default view)
mewt results --all           # Every outcome, not just uncaught
mewt results --format json   # json | sarif | ids | table

# Narrow down (these filters work on both `results` and `print mutants`)
mewt results --target 'src/auth/**'   # Quote globs so the shell does not expand them
mewt results --severity high,medium
mewt results --mutation-types ER,CR
mewt results --status Uncaught        # Uncaught | TestFail | Skipped | Timeout
mewt results --line 42

# Investigate and re-test
mewt print mutant --id [id]              # View the mutated code
mewt test --ids [ids]                    # Re-test specific mutants
mewt test --ids-file uncaught_ids.txt    # Re-test IDs from a file, or '-' for stdin

# Inspect configuration
mewt print config                        # Effective config
mewt print targets                       # Files actually mutated
mewt print mutations --language [lang]   # Mutations and severities for a language
```

Language labels are canonical `family` or `family/dialect` values in mewt 4.x — for example `rust`, `javascript/ts`, `move/sui`, `move/iota`.

---

## What Results Mean

- **Caught/TestFail**: tests detected the mutation (good)
- **Uncaught**: tests did not detect the change. Inspect the code to distinguish a testing gap from an equivalent mutation.
- **Timeout**: tests took too long — inconclusive, not evidence of coverage
- **Skipped**: a less severe mutant was skipped because a more severe mutant on the same line was uncaught

---

## Interpreting Mutation Types

`mewt print mutations --language [lang]` lists every mutation slug, description, and severity for a language, and is authoritative — the operator set grows with each release. What that output does not tell you is what a survivor *means*, which is where prioritization comes from:

| Severity | Representative slugs | What an uncaught mutant tells you |
|----------|---------------------|-----------------------------------|
| High | `ER` (Error Replacement) | Tests tolerate the injected error. Investigate whether the path executes, whether error handling masks the change, and whether assertions check the outcome. |
| Medium | `CR` (Comment Replacement) | Removing the statement does not fail the tests. Check whether its effects matter and whether assertions observe them. |
| Medium | `IF`/`IT` (If False/True), `NR` (Negation Removal) | Tests do not distinguish the changed condition. Both constant replacements surviving can indicate an unexecuted condition or weak assertions on the branch outcomes. |
| Low | Operator shuffles (`AOS`, `COS`, `LOS`, `BOS`, shift/assignment variants), `BL`, `AS`, `LC`, `WF` | Check boundary inputs, arithmetic assertions, and semantic equivalence. The mutation result alone does not establish whether the code executed. |

Severity ranks the *mutation*, not the risk. A low-severity survivor in a fee calculation matters more than a high-severity survivor in a log line — weigh what the mutated code does. Filter with `--severity` to work through the results in priority order.

Attribution

trailofbitstrailofbits
View sourceSee grades on GitHubMore from trailofbits →
SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments (0)

No comments yet. Be the first to comment!

SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Related Skills

Competitor Analysis

This skill provides comprehensive analysis of competitor SEO and GEO strategies, revealing what's working in your market and identifying opportunities to outperform the competition.

1823 votes

Deep Research

Universal deep research agent team. 13-agent pipeline for rigorous academic research on any topic. 8 modes: full research, quick brief, paper review, lit-review, fact-check, three-way literature scan, Socratic guided research dialogue, and systematic review with optional meta-analysis. Covers research question formulation, Socratic mentoring, methodology design, systematic literature search, source verification, cross-source synthesis, risk of bias assessment, meta-analysis, APA 7.0 report co...

502942 votes

Paperclip Distill

Use when an operation issue is a Paperclip cursor-window, distill, or backfill — `operationType: "distill"` or `"backfill"` and the body references a Paperclip source bundle for a project or root issue. Turn raw Paperclip activity into a wiki-insightful project page, decisions log, and history note. This skill exists specifically to replace the stiff, datestamp-heavy templated output that the deterministic distiller produces.

953191 votes

Academic Pipeline

Orchestrator for the full academic research pipeline: research -> write -> integrity check -> review -> revise -> re-review -> re-revise -> final integrity check -> finalize. Coordinates deep-research, academic-paper, and academic-paper-reviewer into a seamless 10-stage workflow with mandatory, coverage-bounded integrity checks, two-stage peer review, and auditable quality-assurance artifacts. Triggers on: academic pipeline, research to paper, full paper workflow, paper pipeline, end-to-end p...

502941 votes

Last30days 2

Research any topic across Reddit, X/Twitter, and the web from the last 30 days. Synthesizes findings into actionable insights or copy-paste prompts.

6511 votes
View all in research →