Skills DirectorySkills Directory
SkillsLearnSecurityCategoriesDocsBlogPro
Sign InSubmit Skill
Skills Directory

Security-tested agent skills for Claude, coding agents, and AI workflows.

Directory

  • Browse Skills
  • All Skills A–Z
  • Claude Skills
  • Claude Code Skills
  • Agent Skills
  • Categories
  • Authors
  • Submit a Skill

Learn

  • Learn Hub
  • Install Claude Skills
  • Write SKILL.md
  • Skills vs MCP
  • Directories Compared

Security

  • Security
  • Methodology
  • Secure Claude Skills
  • Security Badges
  • Chrome Extension
  • Skill Manager

Company

  • About
  • Community
  • Blog
  • API Docs
  • Advertise

2026 Skills Directory. All rights reserved.

ProTermsPrivacyRefunds
Back to skills

Extract

ASecurity

Harvest terms and their definitions from a gathered corpus into terms.jsonl, each with the citation that supports it

3 stars
0 votes
0 copies
1 views
Added 9/19/2026
ai-agentsgobashgit

Works with

cli

Security Analysis

A100/100

Scanned 9/19/2026

$npx -y skills add tony/skills --skill extract --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Extract?

Add the live security badge to your README — it updates automatically with every re-scan.

Security grade badge for Extract
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/tony-extract/badge)](https://www.skillsdirectory.com/skills/tony-extract)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
Files
SKILL.md
---
name: extract
description: "Harvest terms and their definitions from a gathered corpus into terms.jsonl, each with the citation that supports it"
allowed-tools: ["Bash", "Read", "Grep", "Glob", "Write", "Task"]
argument-hint: "[--out=<dir>] [--from=<sources.jsonl>]"
user-invocable: true
disable-model-invocation: true
---


# /scholar:extract

Stage 2. Harvest the terms the corpus uses for its own things, each with a
citation that shows the term in use.

Read `../../references/citation.md` for what a citation must carry and
`../../references/stage-gates.md` for this stage's exit condition.

User arguments: $ARGUMENTS

## What counts as a term

A term is a word the corpus uses for one of its own concepts. In code that is
the names it gives things: type and class names, module and package names,
recurring verbs in function names and commit subjects, the roles in its
configuration, the nouns in its error messages. In prose it is the author's
own vocabulary, taken from `annotate`'s quotation set.

A term is not a word that merely appears. `handler` is a term where the
project distinguishes handlers from services; it is noise where it is just
what someone called a function once.

## Procedure

### 1. Read the source list

Take `sources.jsonl` from `gather`. Harvest only from sources marked
`"read": true`.

### 2. Harvest

For code, the cheapest high-yield passes, all portable:

```console
$ rg -o --no-filename -r '$1' '\b(?:class|struct|interface|type)\s+(\w+)' | sort | uniq -c | sort -rn
```

```console
$ git log --format=%s | rg -o '^\w+' | sort | uniq -c | sort -rn
```

Read the results; do not paste them into the table. Frequency proposes a
candidate, the corpus's own definition confirms it.

### 3. Write terms.jsonl

One row per term as the corpus spells it:

```
{"schema": 1, "term": "adapter", "definition": "wraps a third-party client behind the port interface", "parent": "", "kind": "role", "tier": "structural", "cites": [{"url": "https://github.com/OWNER/REPO/blob/v2.40.0/src/ports.py#L12-L18", "locator": "src/ports.py:12-18", "quote": "every third-party client enters through a port"}]}
```

`parent` stays empty here. `distill` assigns it.

`schema` is the row format's version. It costs one integer now and cannot be
added later without guessing what unversioned rows meant.

`definition` is drawn from the corpus, not composed. Where the corpus never
defines the term, record the definition its usage implies and say so in the
citation's quote by choosing a passage that shows the usage.

## Rules

- Do not invent a term the corpus does not use.
- Do not normalize spellings. `bump` and `update` staying separate is what
  lets `contest` find the synonymy; merging them here destroys the evidence.
- Every row has a non-empty definition and at least one citation with a
  locator, or this stage does not hand off.
- A citation whose quoted text contains the term it supports is circular and
  does not count. Pick a passage that shows the term in use instead.
- Set `tier` to `structural`. Distributional claims are `contest`'s to make,
  and they go in `evidence/metrics.md`, not on a term row.

## Output

Open with a one-line hero (`✓ <n> terms extracted from <n> sources` or
`⚠ Blocked: <reason>`), then exactly these sections:

1. `## Terms` — count by `kind`, and the ten most frequent with their
   definitions.
2. `## Rejected` — candidates that looked like terms and were not, with why.
   This section is the one that shows the harvest was judged rather than
   scraped.
3. `## Gate` — confirmation that every row has a definition and a citation.

End with an `AskUserQuestion` panel offering next steps (for example: run
distill, harvest another source, stop here) — skip the panel only in plan mode.

Attribution

tonytony
View sourceSee grades on GitHubMore from tony →
SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments (0)

No comments yet. Be the first to comment!

SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Related Skills

Caveman

Terse caveman voice: answer first, fluff gone, every technical fact kept. Use for /caveman, "caveman mode", "talk like caveman", "be brief", "less tokens". Stays on until "stop caveman" or "normal mode".

1100021 votes

Hyperplan

Adversarial multi-agent planning skill. Self-orchestrates 5 hostile category members (unspecified-low, unspecified-high, deep, ultrabrain, artistry) via team-mode for ruthless cross-critique debate, distills only the defensible insights, then MANDATORILY hands the distilled insight bundle to the `plan` agent for executable plan formalization. Use when planning needs maximum rigor and surfacing of weak assumptions, blind spots, and over-engineering. Triggers: 'hyperplan', 'hpp', '/hyperplan', ...

698621 votes

Writing Skills

Create and manage Claude Code skills in HASH repository following Anthropic best practices. Use when creating new skills, modifying skill-rules.json, understanding trigger patterns, working with hooks, debugging skill activation, or implementing progressive disclosure. Covers skill structure, YAML frontmatter, trigger types (keywords, intent patterns), UserPromptSubmit hook, and the 500-line rule. Includes validation and debugging with SKILL_DEBUG. Examples include rust-error-stack, cargo-dep...

3931 votes

Mcp Code Execution

Routes multi-tool workflows through MCP servers for large datasets and pipelines. Use when Bash tool overhead is limiting throughput on data-heavy tasks.

3421 votes

catchup

Recovers the conversation and failed tool calls of a previous Codex, Amp, Claude Code, Antigravity, Cline, Copilot CLI, Cursor, DeepSeek Harness, Grok Build, Kimi, OpenCode, Pi Agent, or ZCode session. Use when the user says "catch up", "what did the last session do", "get me up to speed", "I switched agents", asks to recover/summarize a previous session before continuing, or asks to diagnose or report a catchup failure. Do NOT use for the current conversation, git history, or any non-agent log.

741 votes
View all in ai-agents →