Skills DirectorySkills Directory
SkillsLearnSecurityCategoriesDocsBlogPro
Sign InSubmit Skill
Skills Directory

Security-tested agent skills for Claude, coding agents, and AI workflows.

Directory

  • Browse Skills
  • All Skills A–Z
  • Claude Skills
  • Claude Code Skills
  • Agent Skills
  • Categories
  • Authors
  • Submit a Skill

Learn

  • Learn Hub
  • Install Claude Skills
  • Write SKILL.md
  • Skills vs MCP
  • Directories Compared

Security

  • Security
  • Methodology
  • Secure Claude Skills
  • Security Badges
  • Chrome Extension
  • Skill Manager

Company

  • About
  • Community
  • Blog
  • API Docs
  • Advertise

2026 Skills Directory. All rights reserved.

ProTermsPrivacyRefunds
Back to skills

Mcp Server Testing

ASecurity

Prove an MCP server works and that an agent can actually use it. Covers the Inspector, a ten-question evaluation set, and the difference between responding correctly and being usable.

4 stars
0 votes
0 copies
0 views
Added 9/20/2026
ai-agentstypescriptpythongoshellbashtestinggitapidocumentation

Works with

cliapimcp

Security Analysis

A100/100

Scanned 9/20/2026

$npx -y skills add fabioc-aloha/Alex_Skill_Mall --skill mcp-server-testing --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Mcp Server Testing?

Add the live security badge to your README — it updates automatically with every re-scan.

Security grade badge for Mcp Server Testing
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/fabioc-aloha-mcp-server-testing/badge)](https://www.skillsdirectory.com/skills/fabioc-aloha-mcp-server-testing)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
Files
SKILL.md
---
name: mcp-server-testing
description: "Prove an MCP server works and that an agent can actually use it. Covers the Inspector, a ten-question evaluation set, and the difference between responding correctly and being usable."
lastReviewed: 2026-09-13
---

# Test an MCP Server

Two different questions. **Does it respond correctly?** is testing. **Can an agent
accomplish a task with it?** is evaluation. A server can pass the first
completely and fail the second, and the second is the one users feel.

## The Inspector

An interactive client that lists your tools, calls them with arguments you
supply, and shows the raw protocol traffic.

```bash
# TypeScript
npm run build && npx @modelcontextprotocol/inspector ./dist/server.js

# Python
python -m py_compile your_server.py
npx @modelcontextprotocol/inspector -- python your_server.py

# Any executable
npx @modelcontextprotocol/inspector /path/to/server
```

If a tool does not appear in the Inspector, no agent will see it either. Check
this before suspecting anything subtler.

## Pre-Ship Checks

- Every tool has a description saying **when** to use it, not only what it does
- Errors return actionable text with `isError: true`
- I/O is async, and paginated wherever results can grow
- Nothing writes to stdout on a stdio server
- Annotations are declared, especially `destructiveHint`
- Input schemas carry constraints, so bad input fails at the boundary

## Evaluation

Testing proves the server responds. Evaluation proves an agent can *use* it — and
it is the only way to discover that a tool description is technically accurate and
practically useless.

Write ten questions requiring real tool use, then answer each yourself so you hold
ground truth. Each question should be:

- **Independent** — not reliant on another question's answer
- **Read-only** — no destructive operations
- **Multi-step** — several tool calls and some exploration
- **Realistic** — something a person would actually ask
- **Verifiable** — one clear answer, checkable by string comparison
- **Stable** — the answer will not drift next week

```xml
<evaluation>
  <qa_pair>
    <question>Which repository had the most merged pull requests last quarter, and how many?</question>
    <answer>servers, 47</answer>
  </qa_pair>
</evaluation>
```

### Reading the results

A failure is rarely a bug. Work through these in order:

| Symptom | Usual cause |
| --- | --- |
| Agent never calls the right tool | Description says what, not when |
| Agent calls it and gets lost | Response too large, or unstructured |
| Agent gives up after an error | Error text has no recovery path |
| Agent takes ten calls for one answer | Missing a workflow tool over the raw API |

That last row is the signal `mcp-server-design` defers to: add workflow tools once
evaluations show which sequences agents keep repeating.

### Re-run on change

Run the set whenever you change a tool description or schema. A drop in answer
quality is usually a description that got vaguer, not a code regression. Treat the
evaluation set as a test suite and keep it in the repository — see
[test-driven-development](https://github.com/fabioc-aloha/Alex_ACT_ONE) for the
discipline if the project has the ACT runtime installed.

## Testing Under a Host

The Inspector is a clean room. Hosts are not. Before declaring a server done,
register it in a real host and run a task end to end. Differences that only appear
there:

- Tool-name collisions with other installed servers
- Descriptions truncated in the host's picker
- Environment variables resolved differently than in your shell
- Startup time long enough that the host gives up

## Composes With

- [mcp-server-build](../mcp-server-build/SKILL.md) — what you are testing
- [mcp-server-operations](../mcp-server-operations/SKILL.md) — for failures that
  only appear in production

## Would Revise If

- The Inspector is superseded by an official alternative.
- Teams report writing evaluation sets once and never re-running them, which
  would mean this section is documentation rather than practice and needs a
  cheaper harness.

Attribution

fabioc-alohafabioc-aloha
View sourceSee grades on GitHubMore from fabioc-aloha →
SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments (0)

No comments yet. Be the first to comment!

SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Related Skills

Caveman

Terse caveman voice: answer first, fluff gone, every technical fact kept. Use for /caveman, "caveman mode", "talk like caveman", "be brief", "less tokens". Stays on until "stop caveman" or "normal mode".

1100021 votes

Hyperplan

Adversarial multi-agent planning skill. Self-orchestrates 5 hostile category members (unspecified-low, unspecified-high, deep, ultrabrain, artistry) via team-mode for ruthless cross-critique debate, distills only the defensible insights, then MANDATORILY hands the distilled insight bundle to the `plan` agent for executable plan formalization. Use when planning needs maximum rigor and surfacing of weak assumptions, blind spots, and over-engineering. Triggers: 'hyperplan', 'hpp', '/hyperplan', ...

698621 votes

Writing Skills

Create and manage Claude Code skills in HASH repository following Anthropic best practices. Use when creating new skills, modifying skill-rules.json, understanding trigger patterns, working with hooks, debugging skill activation, or implementing progressive disclosure. Covers skill structure, YAML frontmatter, trigger types (keywords, intent patterns), UserPromptSubmit hook, and the 500-line rule. Includes validation and debugging with SKILL_DEBUG. Examples include rust-error-stack, cargo-dep...

3931 votes

Mcp Code Execution

Routes multi-tool workflows through MCP servers for large datasets and pipelines. Use when Bash tool overhead is limiting throughput on data-heavy tasks.

3421 votes

catchup

Recovers the conversation and failed tool calls of a previous Codex, Amp, Claude Code, Antigravity, Cline, Copilot CLI, Cursor, DeepSeek Harness, Grok Build, Kimi, OpenCode, Pi Agent, or ZCode session. Use when the user says "catch up", "what did the last session do", "get me up to speed", "I switched agents", asks to recover/summarize a previous session before continuing, or asks to diagnose or report a catchup failure. Do NOT use for the current conversation, git history, or any non-agent log.

741 votes
View all in ai-agents →