Skip to content
Back to skills

Context Doctor

ASecurity

Audits a repo's agent context — system prompt, CLAUDE.md/AGENTS.md, skills, and references — against context-engineering lessons for Claude 4/5 models (model detected from the running session, or set via --model), then writes claude-optimization.md with a Summary table. Use to right-size an agent setup, cutting over-constraint, redundancy, always-upfront context, and conflicting instructions.

  • 9 stars
  • 0 votes
  • 0 copies
  • 0 views
  • Added October 6, 2026
ai-agentsgoexpresstestinggitsecuritydocumentation

Works with

  • claude code
  • claude desktop
  • cursor
  • cli

Security analysis

A100/100

Pro scans all 3 files and shows the line behind each finding

Scanned October 6, 2026

npx -y skills add sunitghub/canon-skills --skill context-doctor --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Context Doctor?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Context Doctor
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/sunitghub-context-doctor/badge)](https://www.skillsdirectory.com/skills/sunitghub-context-doctor)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: context-doctor
description: Audits a repo's agent context — system prompt, CLAUDE.md/AGENTS.md, skills, and references — against context-engineering lessons for Claude 4/5 models (model detected from the running session, or set via --model), then writes claude-optimization.md with a Summary table. Use to right-size an agent setup, cutting over-constraint, redundancy, always-upfront context, and conflicting instructions.
category: agent-ops
tags: [context, prompt, skills, audit, optimization]
---

# Context Doctor

Static audit of a repository's **agent context** — everything a coding agent loads before it sees a
user prompt: `CLAUDE.md`/`AGENTS.md`, skill and command files, tool descriptions, and referenced
specs/mockups. Rates each of seven lenses, prints a **Summary table**, and writes
`claude-optimization.md` at the repo root. Human-facing name: **Context Doctor**.

The seven lenses come from Anthropic's guidance on context engineering for modern Claude models:
https://claude.com/blog/the-new-rules-of-context-engineering-for-claude-5-generation-models — the
same lessons behind removing ~80% of Claude Code's system prompt. This skill is a portable checkup
against those lessons.

**Self-contained.** It depends on nothing outside this folder — no build tools, no other skills, no
network. Drop it into any repo's `.claude/skills/` or upload it to Claude Desktop and run it.

## When to use

Triggers: "audit my agent setup", "is my CLAUDE.md bloated", "right-size my skills", "check my
context for over-constraint / conflicting instructions", "run context-doctor". Run it periodically,
or after a CLAUDE.md/skills grow large.

Not for: running or testing the target app (static read only), general code review, or a single
sprint diff.

## Target model

Two checkup modes, gating which Claude-5-specific checks run:

- **`5`** (default) — full checkup for a Fable 5/5.1 or Opus 5.5 target: the seven lenses below,
  plus four checks (reasoning-extraction avoidance, effort-default guidance, checkpoint/pause
  discipline, progress-claim grounding). A repo targeting Opus 5 also runs in `5`, but check 10 is
  not raised for it and check 11 uses the `high` default.
- **`4`** — the seven lenses only, plus the two checks that are model-agnostic (checkpoint/pause
  discipline, progress-claim grounding). Skip reasoning-extraction avoidance and effort-default
  guidance — those failure modes are specific to Claude 5-generation models and would be false
  positives for a repo targeting Opus 4.8.

Detect the default from the running session's own model identity (stated in the system prompt).
Override with `--model 4` or `--model 5` when the repo's target differs from the session model —
e.g. auditing on Sonnet 5 a repo whose skills are written for an Opus 4.8 production deploy. State
which mode was used, and how it was determined (detected vs. `--model` override), in the report
header.

## Operating constraints

- **Static analysis only.** Read the context files. Never run the repo, execute skills, or call a
  model. No dynamic probing.
- **Judgement, not a checklist score.** Report a per-lens status and one overall verdict — **never a
  numeric score or percentage**. A number hides which lens is weak.
- **Evidence, not theory.** Cite `file:line` (or `file` + a short quote) for every finding. Flag a
  lens `action`/`advisory` only with a concrete instance, not a hypothetical.
- **Repo-agnostic + graceful.** If an artifact is absent (`CLAUDE.md`, `.claude/skills/`, etc.), say
  so and continue — an absent file is a valid result, not a finding.
- **Read, don't rewrite.** This skill diagnoses and recommends. It does not edit the audited files;
  it only writes the one report.

## What to inspect

Gather the context artifacts that exist in the target repo (skip any that are absent, note which):

- `CLAUDE.md` (repo root and any nested), `AGENTS.md`, `.cursorrules`/other agent-instruction files
- `.claude/skills/` (or `skills/`) — each `SKILL.md`, plus any `reference/`/`gates/` sub-files
- Tool/command definitions and their descriptions
- `@`-imported or referenced standards/config injected into every session
- Referenced specs, plans, mockups (are they prose, or code/HTML/tests?)

Line counts are a proxy for context weight, not exact tokens.

## The seven lenses

Rate each: **aligned** (follows the lesson), **advisory** (minor drift, worth trimming), or
**action** (clear instance to fix). Cite evidence.

1. **Rules → judgement.** Blanket prohibitions/mandates a capable model handles via judgment.
   - Check for absolute rules that are wrong in some cases: "never write comments", "always
     do X", rigid formatting dictates. Prefer judgment-framed guidance ("match the surrounding
     code's comment density, naming, and idiom").
   - `action` when a blanket rule would produce wrong behavior for a reasonable subset of tasks.

2. **Examples → interface design.** Over-reliance on usage examples where an expressive interface
   would guide better.
   - Check tool/command definitions: do they lean on long "here's how to call it" examples, or do
     clear parameter names, enums, and defaults make correct use obvious?
   - `advisory` when examples substitute for a self-describing interface.

3. **Upfront → progressive disclosure.** Context injected on every session that is only
   conditionally needed.
   - Check for always-loaded content used by a minority of tasks (verification/review steps, rare
     workflows, deep references). Recommend moving it behind on-demand skills/reference files loaded
     when needed. Note oversized always-on files (a long root `CLAUDE.md`, a monolithic SKILL.md).
   - `action` when a large block is always injected but rarely used.

4. **Repeat yourself → simple descriptions.** The same instruction duplicated across places.
   - Check for guidance repeated in the system prompt/`CLAUDE.md` *and* a tool/skill description, or
     the same rule copied across files. Recommend one owner; put tool usage in the tool description.
   - `advisory`/`action` per how much duplication and drift risk exists.

5. **Memory: manual → durable.** How cross-session knowledge is captured.
   - Check for heavy manual "save this to memory" instructions. Note that modern harnesses can
     auto-capture relevant memory. Portable, repo-native memory (decision logs, handoff notes) is a
     legitimate deliberate choice — flag only redundant manual bookkeeping, not durable records.
   - Usually `advisory`.

6. **Simple specs → rich references.** Fidelity of the references the agent works from.
   - Check whether specs/designs are prose or screenshots where a higher-fidelity reference exists:
     a code file to port, a test suite as the spec, an HTML mockup instead of a description or
     screenshot, or a rubric a verifier can check against. Recommend code/HTML/test references.
   - `advisory`/`action` when a prose/screenshot reference could be a code/HTML/test artifact.

7. **Conflicting instructions.** Contradictory directives across the loaded context.
   - Cross-read the artifacts for clashes (e.g. "leave documentation as appropriate" vs "DO NOT add
     comments"; "keep it minimal" vs "be thorough"). Contradictions force the model to spend
     reasoning reconciling them. Quote both sides.
   - `action` for any direct contradiction; `advisory` for tension worth clarifying.

## Model-agnostic checks (both `4` and `5`)

8. **Checkpoint/pause discipline.** Instructions that make Claude stop or ask permission more than
   the task needs.
   - Check for guidance that would block on "Want me to…?"/"Shall I…?" for reversible, in-scope
     actions, or that omits when pausing is actually warranted (destructive/irreversible actions,
     real scope changes, input only the user can provide).
   - `action` when a skill/CLAUDE.md lacks any checkpoint guidance for a long-running or autonomous
     workflow; `advisory` when guidance exists but is vague.

9. **Progress-claim grounding.** Long-run status reporting that isn't tied to verifiable evidence.
   - Check whether workflows that report progress (multi-step skills, autonomous loops) instruct
     grounding each claim in a tool result from the session, and stating explicitly what wasn't
     verified.
   - `advisory` unless the repo has evidence of prior fabricated status reports, then `action`.

## Claude 5-specific checks (`5` only)

Which model the repo targets decides the details of both checks. Take it from the model ids the
repo's own skills/agents/settings name; if none, from the running session's model; if still unclear,
report both models' guidance as `advisory` rather than picking one. (Sources: Anthropic's
[Prompting Claude Opus 5.5](https://platform.claude.com/docs/en/build-with-claude/prompt-engineering/prompting-claude-opus-5-5)
and [Effort](https://platform.claude.com/docs/en/build-with-claude/effort) pages.)

10. **Reasoning-extraction avoidance.** Instructions that ask Claude to echo, transcribe, or explain
    its internal reasoning as response text.
    - Check skills/CLAUDE.md for "explain your reasoning", "show your work", or similar asked of the
      *response* (not `thinking` blocks). This can trigger the `reasoning_extraction` refusal
      category on **Fable 5 and Opus 5.5** (new on Opus 5.5 relative to Opus 5). Fix: drop the
      instruction and read the reasoning from summarized thinking blocks instead. Server-side
      fallback does not retry a `reasoning_extraction` decline on either model — it comes back to the
      caller, so another model does not silently take over.
    - `action` when found — do not raise this check in `4` mode, or for a repo targeting Opus 5
      (which lacks the category); the Anthropic pages don't document it for earlier models, so
      treat it as not applicable there rather than asserting it.

11. **Effort-default guidance.** Whether effort-level usage matches the target model's cost/latency
    curve. The defaults differ by model:
    - **Fable 5 / 5.1:** default `high`; `xhigh` for the most capability-sensitive workloads;
      `medium`/`low` for routine work.
    - **Opus 5:** default `high`, like Fable.
    - **Opus 5.5:** default `medium` (Opus 5 defaults to `high`, so a carried-over setting runs a
      level off). Set effort explicitly, sweep levels against the repo's
      own evals, and reserve `xhigh`/`max` for work with a measured quality gain.
    - Check for blanket `xhigh`/`max` on routine work, or no effort guidance at all on a repo doing
      capability-sensitive work, and for an Opus 5.5 repo that assumes `high` is the default.
    - `advisory` — these are tuning recommendations, not a correctness bug; do not raise this check
      in `4` mode, since Opus 4.8's effort/quality tradeoff differs. Judgement-heavy review gates
      (canon's own reviewer/evaluator) can legitimately stay at `high`: the source itself says to
      test before lowering effort.

## Summary and verdict

Open the report with a Summary table — one row per lens:

| Lens | Status | Evidence | Recommendation |
|---|---|---|---|

Then an overall posture verdict:

| Verdict | Criteria |
|---|---|
| **lean** | No `action` lenses; at most minor `advisory` notes. Context is well right-sized. |
| **trim** | One or more `action` lenses, each with a concrete, scoped fix. Worth a cleanup pass. |
| **overloaded** | Multiple `action` lenses or a structural problem (large always-on context, several conflicts) needing a deliberate restructure. |

No numeric score. Any lens rated `action` forces at least **trim**.

## Report

Print the Summary table inline. Then ask: `Write claude-optimization.md to the repo root? (y to confirm)`.
Do not write without `y`. On confirmation, write `claude-optimization.md` at the repo root (overwrite
— it is a point-in-time snapshot, not a log):

```
context-doctor run: MM-DD-YYYY hh:mm
Audited against: Claude <4|5> context-engineering guidance (<detected from session | --model override>)

## Context optimization: <repo-name>
Scope: <artifacts inspected; which were absent>

| Lens | Status | Evidence | Recommendation |
|---|---|---|---|
| Rules → judgement | <status> | <file:line or quote> | <one-line fix> |
| Examples → interfaces | ... | ... | ... |
| Upfront → progressive disclosure | ... | ... | ... |
| Repeat → simple descriptions | ... | ... | ... |
| Memory: manual → durable | ... | ... | ... |
| Specs → rich references | ... | ... | ... |
| Conflicting instructions | ... | ... | ... |
| Checkpoint/pause discipline | ... | ... | ... |
| Progress-claim grounding | ... | ... | ... |
| Reasoning-extraction avoidance (5 only) | ... | ... | ... |
| Effort-default guidance (5 only) | ... | ... | ... |

Omit the last two rows entirely in `4` mode — don't print them as "n/a".

### Details
<one short paragraph per lens rated advisory/action, each with file:line evidence>

### Not inspected
- <artifact absent — why it was skipped>

context-doctor verdict: lean | trim | overloaded
```

The final `context-doctor verdict:` line is required.

## Gotchas

- **No context artifacts found?** If the repo has no `CLAUDE.md`/`AGENTS.md`/skills/agent config,
  say so and stop with a limited-scope note — do not invent findings. A repo with no agent context
  is a valid, clean result.
- **Durable memory is not clutter.** A decision log or handoff file is deliberate cross-session
  memory, not the manual-bookkeeping the memory lens flags. Don't recommend deleting durable records.
- **Don't over-apply "rules → judgement".** Constraints that guard genuinely dangerous or
  irreversible actions (destructive commands, security boundaries, mandatory review gates) should
  stay explicit — the lesson is to relax *advisory* over-constraint, not safety-critical rules.
- **No numeric score.** If tempted to write "7/10" or a percentage, stop — use per-lens status plus
  the lean/trim/overloaded verdict.
- **Optional canon companion.** In a repo that uses canon, `context-check` gives a deeper always-on
  *budget* audit (line-by-line size/redundancy). context-doctor does not require it and never calls
  it — mention it only as a follow-up if present.

Files in this skill

  • README.md2.5 KB
  • SKILL.md13.9 KB
  • evals/evals.json5 KB

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…