Checks in about a minute whether an MCP server's metadata survives the clients it will run in (Claude Code, claude.ai, Cowork, ChatGPT Chat and Work, Codex): reads tools/list and the server instructions over HTTP or stdio, and scores each tool against each client's measured limits on descriptions, server instructions, input schemas and output schemas. Invoke when asked to check, health-check, lint or score an MCP server, to ask "will this server work in ChatGPT / Codex / claude.ai", "why does...
Scanned 9/28/2026
Install to Claude Code
npx -y skills add DaveGold/mcp-metadata-demo --skill mcp-compat-check --agent claude-codeInstalls into .claude/skills of the current project.
Are you the author of Mcp Compat Check?
Add the live security badge to your README — it updates automatically with every re-scan.
[](https://www.skillsdirectory.com/skills/davegold-mcp-compat-check)More formats (shields.io, HTML) on the badges page.
---
name: mcp-compat-check
description: >-
Checks in about a minute whether an MCP server's metadata survives the clients it will run in
(Claude Code, claude.ai, Cowork, ChatGPT Chat and Work, Codex): reads tools/list and the server
instructions over HTTP or stdio, and scores each tool against each client's measured limits on
descriptions, server instructions, input schemas and output schemas. Invoke when asked to check,
health-check, lint or score an MCP server, to ask "will this server work in ChatGPT / Codex /
claude.ai", "why does the agent ignore my description or parameter docs", or to compare public
servers. Hands fixes to the rich-domain-mcp-server skill.
metadata:
version: 1.0.0
---
# mcp-compat-check — does the metadata reach the model on every client?
A server can ship perfect metadata and still not deliver it, because each client cuts or drops
different surfaces. This check compares what a server ships with what each client delivered when
it was measured (2026-09-27; evidence: `rich-domain-mcp-server` → `references/evidence.md`, tag
HD). It is static: no model and no tool calls, so it runs in seconds and gives the same answer
every time.
## Run it
```bash
node .claude/skills/mcp-compat-check/check.mjs https://example.com/mcp
node .claude/skills/mcp-compat-check/check.mjs --header "Authorization: Bearer $TOKEN" https://example.com/mcp
node .claude/skills/mcp-compat-check/check.mjs --stdio -- node dist/stdio.js
node .claude/skills/mcp-compat-check/check.mjs --json https://example.com/mcp # for CI or diffing
node .claude/skills/mcp-compat-check/check.mjs --oauth https://mcp.example.com/mcp # OAuth servers (needs @modelcontextprotocol/sdk)
```
Never put a token in the command yourself: ask the user to export it as an environment variable,
and pass the variable. For OAuth servers use `--oauth`: it opens the user's browser, the user signs
in and approves, and the token lives only in the checking process. Run one OAuth server at a time,
and tell the user which service the next tab is for.
## What it checks, per client column
| surface | Claude only (Code, claude.ai, Cowork) | ChatGPT / Codex only | all clients |
|---|---|---|---|
| tool description | ≤ 2,048 (Claude Code cuts there) | no cut measured | ≤ 2,048 |
| server instructions | ≤ 2,048, and nothing required (claude.ai chat gets none) | ≤ 512, and nothing required (Work gets none) | ≤ 512, nothing required |
| input schema | whole | ≤ 5,000 serialized, or every describe is dropped (bisected on Codex; ChatGPT Work consistent, not bisected) | ≤ 5,000 |
| output schema | never delivered: meaning there is review-only | delivered on Codex and Work | review-only |
It also flags: guidance that lives mostly in the server instructions, tools with a first sentence
too short to be found by clients that load tools by search, missing annotations, and the size of
`tools/list`, which is re-sent every turn.
## Reading the report
The **Verdict** table has one column per client group and one row per surface:
| row | what it counts |
|---|---|
| **overall** | the worst surface in that column, which surfaces fail, and how many tools are affected |
| tool description | tools whose description is longer than the column's cut |
| server instructions | their length against the column's budget, and whether the guidance lives there; within budget is green, with a note that some clients deliver none |
| input schema | tools whose serialized schema is over 5,000, so Codex and ChatGPT Work drop every describe |
| output schema | tools that ship one; on the Claude clients it never arrives, so meaning found only there is flagged |
- 🟢 **portable**: within every measured limit for that column.
- 🟡 **host-dependent**: it arrives on some clients, or is never needed; review it. Output-schema
meaning is fine if the same meaning travels in the field names or the response, which a static
check cannot see.
- 🔴 **not portable**: the server relies on a surface or size that some measured client in that
column does not deliver.
**A portability finding, not a quality rating.** A protocol-correct, well-built server can be red:
the protocol lets it ship metadata that some clients never pass on. The report's last section lists
the rule behind every cell, so anyone can recompute the matrix.
Start with the "Start here" list. For each fix, use the `rich-domain-mcp-server` skill: its
*Budgets by target client* table has the rule and its evidence, and its REQUIRED guidance-tool
pattern holds detail that does not fit a budget.
## Limits
- One day's client versions. Re-measure before relying on a number: the quote probe and the Codex
request trace are in `rich-domain-mcp-server` → `references/delivery.md`.
- It does not call tools, so it cannot see the response, where most meaning should live.
- A green report means the metadata arrives. It says nothing about whether the metadata is right.
Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.
No comments yet. Be the first to comment!