Use when the user wants to audit, score, grade, improve, or converge a repository's documentation system as an AI-agent context source — `measure` (reproducible OK/WATCH/FAIL signals: dead links, orphans, freshness, entry-file token cost) plus `judge` (LLM star ratings for completeness, correctness, consistency), with an improvement loop that fixes docs until every measure signal meets its target. Covers docs quality audit, documentation health check, dead-link/orphan/staleness checks, doc co...
Installs into .claude/skills of the current project.
Are you the author of Docgrad?
Add the live security badge to your README. It updates with every re-scan.
[](https://www.skillsdirectory.com/skills/redtear1115-docgrad)
---
name: docgrad
description: Use when the user wants to audit, score, grade, improve, or converge a repository's documentation system as an AI-agent context source — `measure` (reproducible OK/WATCH/FAIL signals: dead links, orphans, freshness, entry-file token cost) plus `judge` (LLM star ratings for completeness, correctness, consistency), with an improvement loop that fixes docs until every measure signal meets its target. Covers docs quality audit, documentation health check, dead-link/orphan/staleness checks, doc convergence. 中文關鍵字:文件評分、文件健檢、文件收斂、docs 評比、文件品質、死鏈檢查、文件過期。Not for prose style linting, SKILL.md auditing, or code review.
argument-hint: "[init · measure · judge · audit · improve · loop · report]"
license: MIT
---
> **Last updated:** 2026-09-17
Grade and converge a repo's documentation system (the docs directory plus the root instruction
files) as an **AI agent context source**, in two layers: `measure` — four scripts, reproducible
OK/WATCH/FAIL signals (dead links, orphans, freshness, entry-file token cost) — and `judge` — an
LLM star rating for the three dimensions a script can't check (completeness, correctness,
consistency). `loop` fixes docs round by round until every `measure` signal meets its target; it
never fixes for a judge star. It does not lint prose style, does not review code, and does not touch CI.
Prerequisite: the documentation is a **local markdown file tree** and `.docgrad.yml` can be written
to the target repo root; wikis and remote doc sources are not supported (boundaries and workarounds
in [docs/design.md](../../docs/design.md) §Positioning and boundaries).
`SKILL_DIR` = the directory this file lives in (the relative root for `scripts` and `reference`).
## Routing
| User input | Action |
|---|---|
| `/docgrad` (no argument) | Print this table, plus a one-paragraph statement of the two layers: **measure** is reproducible, has targets, and is what CI/`loop` gate on; **judge** is the model's stars, not comparable across rounds, and has no targets. Do nothing else |
| `init` | Read [reference/init.md](reference/init.md) and follow it |
| `measure` | Run [reference/measure.md](reference/measure.md) (the four scripts): report only, changes no files. No rubric.md needed |
| `measure <scope>` | Scoped: limit to a directory, glob, or topic. Still report-only — see judge.md §Scoped audit |
| `judge` | Read [reference/rubric.md](reference/rubric.md), run [reference/judge.md](reference/judge.md) (needs this round's measure.md output — see judge.md's own precondition): report only, changes no files |
| `judge <scope>` / `judge --dim <dimension>` | Scoped judge: limit to a directory, glob, or topic, or rate a single dimension. Still report-only, and it **never writes to `.docgrad/`** — see judge.md §Scoped audit |
| `audit` | **Deprecated alias** (kept for 1.x muscle memory): runs `measure`; runs `judge` too only when passed `--judge`. Report only either way — `audit` never wrote to `.docgrad/`, before or after this alias |
| `audit <scope>` / `audit --dim <dimension>` | Scoped form of the alias. A dimension can only be judged, so `--dim` implies `--judge`: `audit --dim <d>` behaves as `measure` + `judge --dim <d>`. Still report-only; see judge.md §Scoped audit |
| `improve` | Run [reference/measure.md](reference/measure.md) every round; run rubric.md + [reference/judge.md](reference/judge.md) only when passed `--judge`. Run one round per [reference/improve.md](reference/improve.md) |
| `loop` | Same as improve, repeated until a stop condition; `--judge` is a per-invocation flag, not a mode — pass it every time a round should also rate |
| `report` | Read the target repo's `.docgrad/scorecard-latest.md` and reprint it, plus a per-round score trend drawn from `.docgrad/history.jsonl`. If the files do not exist, tell the user to run improve/loop first (a plain audit is report-only and writes nothing). **Check for branch divergence first**: `git rev-list --count HEAD..docgrad/converge` (when that branch exists) > 0 → put a warning at the top of the report: "the converge branch is N commits ahead of this branch, the trend below may be incomplete". The report header always states its data source as `<branch> @ <short-sha>`. **Legacy rows** (no `schema`) are shown as a separate "1.x rounds" block and are never joined to schema-2 rows: draw one break before the first schema-2 row, "v2 changed what is measured; rows above are 1.x and cannot be compared". Inside the legacy block the 1.x rules apply unchanged, kept as the legacy rules: (a) `rubric_hash` differs from the previous round → draw a break line noting "the ruler changed here, scores before and after cannot be compared directly"; (b) `corpus_hash` differs from the previous round → draw a break line noting "the corpus scope changed here: `files_total`, `claims_total`, the freshness denominator and the pollution denominator all moved, so scores either side cannot be compared"; (c) the measure fingerprint — read as `measure_hash`, or as `thresholds_hash` when that is the name present — differs from the previous round → break, same wording as (a); (d) the judge fingerprint — read as `judge_hash`, or as `judgement_hash` when that is the name present — differs from the previous round → break, same wording as (a); (e) rounds with no `economy` key are from the five-dimension era before v1.0.0 → draw `—` for that dimension and note "the rounds below are five-dimension; targets-met and overall scores cannot be compared with newer rounds"; (f) a field missing from a legacy row is unknown and draws no break — this covers every legacy shape: no fingerprints at all, only some fingerprints, old names, and new names without `schema` (rows written between v2.0.0 E1 and this change). This is also where rubric.md §Version history's "report has to map them" is honoured: the name mapping in (c) and (d) is that mapping. **Schema-2 row violations** are each printed as "row N violates history schema 2: <what>": a missing `docgrad` object; a missing `measure` object; or any of these required `docgrad` keys missing — `version`, `measure_hash`, `judge_hash`, `corpus_hash` (the key set `docgradMeta()` emits today). The required-key check applies only when the `docgrad` object is present — a missing object is one problem, not five. Each violating row gets exactly one line listing all of its problems, separated by `; `, for example "row 2 violates history schema 2: missing docgrad" or "row 4 violates history schema 2: missing corpus_hash; missing version". A key not in the list is ignored. In particular, a schema-2 row written before v2.0.0 E4b may carry `rubric_hash`; it is ignored, since `judge_hash` now covers rubric.md. The violating row is still shown, but no break decision is inferred from it. **Comparison baseline**: every schema-2 break decision compares a row with the previous schema-2 row that is not a violation — a violating row is skipped as a baseline. A key that is present with value `null` is a value, not a violation (`corpus_hash` is null without a config, and `version` can be null per lib.mjs `docgradMeta`); null compares equal only to null. **Measure trend**: one series per measure id, number first and verdict beside it; break the measure trend when `docgrad.measure_hash` or `docgrad.corpus_hash` differs from the baseline, naming which one moved. **Judge series** (optional to draw) is labelled "judge — not comparable across rounds", and never averaged or summed; break it when `docgrad.judge_hash` differs from the baseline — a judge-series break alone does not break the measure trend. **No overall score is ever computed.** When `.docgrad/ledger.jsonl` exists, also report cumulative coverage and which claims are still `fail` or `stale` |
## Blockers (non-skippable)
1. The target repo has no `.docgrad.yml` → every command except `init` must first redirect to
`/docgrad init`.
2. Before running `judge` — directly, via `audit --judge`, or via `improve --judge`/`loop --judge` —
you must read [reference/rubric.md](reference/rubric.md); the star anchors may not be invented or
relaxed. Before rating **consistency** you must also read [reference/placement.md](reference/placement.md) —
the rules for judging placement and duplication live there. `measure` (directly, or as part of
`audit`/`improve`/`loop` without `--judge`) needs none of this.
3. improve/loop commit only on the `docgrad/converge` branch, and never modify the target repo's CI
configuration.
## Scripts
Five dependency-free Node (≥18) scripts. They read the target repo's `.docgrad.yml`, write JSON to
stdout (consume it in full — do not truncate it through `head` or `grep`), and report errors on
stderr with a non-zero exit:
```bash
node "$SKILL_DIR/scripts/inventory.mjs" --root . # inventory/tokens/fixed cost/pollution surface/out_of_scope/untracked files/claim candidates/section structure
node "$SKILL_DIR/scripts/links.mjs" --root . # dead links/broken anchors/orphans/reachable ratio
node "$SKILL_DIR/scripts/freshness.mjs" --root . # date-signal coverage/git comparison (convention may be multi-valued)
node "$SKILL_DIR/scripts/coverage.mjs" --root . # coverage drift/undocumented areas
node "$SKILL_DIR/scripts/retrieval.mjs" --root . # traceability/marginal cost (when scenarios is set)
```
Shared flags: `--config <file>` (when the config file is not at the root), `--include <glob>`
(limits the scope for a scoped audit; repeatable or comma-separated. A pattern that matches no file
in the final included set is an error — `inventory.mjs`, `links.mjs` and `freshness.mjs` exit
non-zero, naming every unmatched pattern and why (#120); it no longer silently scopes to an empty
corpus. `coverage.mjs` and `retrieval.mjs` **accept it and deliberately ignore it** — they exit 0,
report `scope: null`, and each explains in its `note` why narrowing the scope would misjudge its own
measurement. They do not reject it, so a scoped run against them does not fail; it silently measures
the full corpus, which is why the note matters),
`--locate-ledger <path>` (path to a claim ledger, #63 — only `inventory.mjs` acts on it, emitting a `locate_ledger` block that says where each ledgered claim sits in this round's corpus, uncapped and read off the unfiltered population; the other four accept it and report it as a no-op in their own `note`),
`--exclude-ledger <path>` (path to `.docgrad/ledger.jsonl`, #54 — only `inventory.mjs` acts on it,
filtering already-ledgered candidates out of `claim_candidates` before `claim_candidates_cap` is
applied; the other four scripts accept it and report it as a no-op in their own `note`).
The `scope` field in each script's output is the scope the report must state.