Turn session and delivery evidence into reusable lessons, so each mistake is only paid for once. Use when the user asks for a retrospective, wants to mine sessions or outcomes, asks what should become a skill or where tokens were wasted, or after a campaign, incident, or review gate needed many rounds.
Scanned 9/20/2026
Install to Claude Code
npx -y skills add askrubberduck/skills --skill duck-learn --agent claude-codeInstalls into .claude/skills of the current project.
Are you the author of Duck Learn?
Add the live security badge to your README — it updates automatically with every re-scan.
[](https://www.skillsdirectory.com/skills/askrubberduck-duck-learn)More formats (shields.io, HTML) on the badges page.
---
name: duck-learn
description: Turn session and delivery evidence into reusable lessons, so each mistake is only paid for once. Use when the user asks for a retrospective, wants to mine sessions or outcomes, asks what should become a skill or where tokens were wasted, or after a campaign, incident, or review gate needed many rounds.
---
# Duck Learn
The feedback loop: evidence from past work becomes durable updates — a skill, a memory, a rule —
or gets consciously discarded. Lessons that live only in a chat transcript are lessons lost.
Follow the user’s language unless they ask otherwise. Keep commands, paths, identifiers,
quoted errors and machine-readable verdicts unchanged; the duck asks for evidence in any language.
## Recipe
1. **Gather evidence, don't reminisce.**
- **Sources.** Transcript-store locations — Claude `~/.claude/projects/<dir>/*.jsonl`, Codex
`$CODEX_HOME/sessions` and `archived_sessions` — are the hosts whose stores are known, not the
whole set: another host has its store located before the mine, or the result is partial and
says so. Alongside them, recorded outcomes, review trajectories, and token stats. Missing or
malformed store? Say so and mark the result partial.
- **Count.** Extract counts: repeated directives, repeated failures, repeated tool patterns.
Count independent owner tasks, not files: fold every derived log — subagents, retries,
forwarded copies — into its parent, by event timestamp rather than file mtime, roots and
delegated logs reported apart. Generated prompts, task notifications, and tool results are
tool evidence, never owner directives.
- **Never raw.** Big transcripts are mined by script or subagent, never read raw into the main
context, and no raw prompt text goes into durable output.
2. **Classify each candidate lesson** by its durable home — one authoritative home per lesson:
- Repeatable multi-step workflow **the user asks for in words** → a **skill** (new, or a section
of an existing one — prefer extending; a new skill is a cost).
- Behavior that must fire on **repo state** rather than phrasing — a campaign left open, a gate
pending, a stale base — → the checked-in instructions doc. **A skill description matches words;
it cannot see state.**
- Fact, preference, or project state → **memory**.
- A **defect class** the doer repeated → the repo's defect ledger, which `duck-proof` reads
before every pass. Classes compound; instances do not.
- Rule that must bind every turn → the checked-in instructions doc (CLAUDE.md/AGENTS.md).
- One-off, derivable, or already recorded → **discard, say so**.
3. **Evidence bar**: 2+ independent occurrences or an explicit owner directive → build it.
One occurrence → park it as a note in the nearest existing home, not a new artifact.
4. **Apply the updates** — write the skill/memory/rule edit now, not a recommendation to write it.
While in each home, delete what the new lesson supersedes; stale guidance is worse than none.
Policy files are the limit, and the limit is authority rather than effort: an edit to a skill, a
gate, or an instruction file is trust-touching, so it travels the same plan and review path as
any other change to them — derive it, write it up, hand it on, never land it unreviewed.
5. **Close the loop**: procedural guidance (a skill, a workflow rule) gets one rep before it's
trusted. Reserve one occurrence as a holdout BEFORE deriving — at exactly two, derive from the
other — state the expected outcome, then run the guidance against that holdout and attack the
result with `duck-proof` discipline; a failed rep sends the guidance back to draft. Check
observable behavior, not whether the agent repeats the new rule: execute a counterexample,
verify final artifacts, and record the candidate guidance, prompt, oracle and result. Pair
opposite user preferences when testing agreement bias. Use an old-guidance baseline before
claiming improvement; one successful rep proves only that case.
Directive-derived guidance has no occurrence to reserve — it stays draft until its first real
occurrence, which serves as its holdout rep. Memory entries instead record their source
occurrence. Guidance that has never fired is a draft, not a lesson.
## Common mistakes
- Saving what the repo already records (git history, code structure) — memory duplicating the repo
rots; link, don't copy.
- Mining only failures — validated approaches that WORKED are equally worth encoding (with their
evidence), or they'll be re-derived at full cost next time.
Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.
No comments yet. Be the first to comment!