Use when writing, refining, or reviewing any prompt or agent instruction that drives real work - agent/subagent prompts, spec/plan/rewrite prompts, interview or interrogation prompts (including designing the questions an agent should ask), meta-prompts, system prompts, and other model-facing instruction payloads such as tool descriptions, eval-judge rubrics, and grading criteria. Trigger whenever the user says "improve this prompt", "write a prompt for X", "refine my instructions", "why does ...
Installs into .claude/skills of the current project.
Are you the author of Refine Prompt?
Add the live security badge to your README. It updates with every re-scan.
[](https://www.skillsdirectory.com/skills/zyeri-refine-prompt)
---
name: refine-prompt
description: >-
Use when writing, refining, or reviewing any prompt or agent
instruction that drives real work - agent/subagent prompts, spec/plan/rewrite prompts,
interview or interrogation prompts (including designing the questions an agent should
ask), meta-prompts, system prompts, and other model-facing instruction payloads such as
tool descriptions, eval-judge rubrics, and grading criteria. Trigger whenever the user
says "improve this prompt", "write a prompt for X", "refine
my instructions", "why does the model ignore this instruction", "the agent keeps picking
the wrong tool", or hands over a prompt/spec to tighten - even if they never say the
word "prompt". The point is to replace vague exhortations ("be thorough", "ask lots of
questions") with mechanisms a model will actually execute. NOT for: shell command
prompts, permission dialogs, UI copy/microcopy, UserPromptSubmit hooks, commit messages,
prose or document editing, human-facing interview or survey design, or merely
translating an existing prompt.
---
# Refine Prompt
A prompt is a program written for a model. Refining one is not editing for style -
it is predicting where the draft will make the model do something generic or lazy,
and replacing that spot with a mechanism that forces the behavior you actually want.
The governing idea: **vague words optimize for the wrong thing.** "Ask many questions"
optimizes question _count_; "cover every fork" optimizes _coverage_. Same intent, and
only one of them is executable.
## Scope gate - pick the depth before you start
| Situation | What to do |
| ---------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------- |
| Prompt drives substantial work AND user wants it sharpened | **Full method** (Gather -> Diagnose -> Rewrite -> Verify) |
| Small ask / quick tweak | **Inline substitutions only** - apply the smell table, skip artifact reading, batched interrogation, and deliverables |
| Diagnose flags nothing - no row predicts filler | **Null verdict** - one line saying no material change is needed, naming what you checked. Do not manufacture a rewrite. |
| Shell prompt, permission prompt, UI dialog, UPS hook | **Not this skill** - leave it alone |
The scope gate matters because the full method has real cost (reading artifacts,
building a simulation table). Spending it on a one-line tweak is the same mistake as
"be thorough" - effort pointed at the wrong thing.
## The method: Gather -> Diagnose -> Rewrite -> Verify
### 1. Gather
- **Read the referenced artifact first** - when one exists, reading is proportionate,
and the user hasn't declined. Extract the real fork points, latent bugs, and honest
constraints, then _enumerate them by name_ in the prompt. A named decision can't be
silently skipped; "leave no stone unturned" is an exhortation, "resolve these 7
decisions" hands over the actual stones. No artifact / too large / user declined ->
enumerate assumptions in the prompt instead.
- **Front-load context.** Put a one-paragraph summary of the target artifact inside the
prompt itself, so it survives being pasted into a fresh session and the model orients
before its first tool call.
- **Preserve the user's intent skeleton.** Note every goal in the original. The refined
prompt is the original _compiled_, never replaced - each goal must map to a mechanism,
and no new scope gets invented.
### 2. Diagnose - simulate, don't judge
Build a table, **one row per thing the model must act on** - merge rows sharing a fix,
cap it at **12**, and note any remainder in one line:
| Instruction (as written) | What the model will literally do | Verdict |
| ------------------------ | -------------------------------- | ------- |
Fill the middle column by role-playing a literal, slightly-lazy model - not the ideal
reader you hope for. If a row predicts generic behavior (filler clarifying questions,
restating the goal back, hedged "it depends" non-answers, offloading reading onto the
user), mark it for rewrite. You are cataloguing _predicted failures_, not stylistic
preferences - this is what gives the Verify step something concrete to check.
Consult the **smell table** below to name each failure and its fix.
### 3. Rewrite - swap smells for mechanisms
Rewrite every flagged row using the matching mechanism from the smell table. Keep the
prompt lean: a mechanism that doesn't change predicted behavior is dead weight, cut it.
### 4. Verify - re-simulate the rewrite
Re-run the simulation table against the _refined_ prompt. Iterate until no row predicts
filler, **capped at 3 rounds**; at the cap, ship it and list the rows still predicting
filler as stated known-weak spots. A rewrite is an unverified hypothesis until you do
this. Then deliver:
1. The refined prompt.
2. A short **change table**: `original phrase -> predicted failure -> mechanism applied`.
This makes the refinement itself checkable by the user, rather than a black-box "trust me".
If Diagnose flagged nothing, deliver the null verdict instead - an empty change table
means there was nothing to compile.
## Smell table - the core substitutions
The recurring failure patterns and what to replace them with. This is also the entire
toolkit for a **quick tweak** (scope gate row 2).
| Smell in the draft | Why it fails | Mechanism to replace it with |
| ----------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Quantity words: "many / all / thorough / as much as possible" | Optimizes the count, not coverage; no termination = no done-state | Themed batches **sized to the actual fork count** (2 forks -> batch of 2), an explicit **stop condition**, and a **round cap**. At the cap, record leftovers as stated assumptions. |
| "Ask the user questions" with no shape | Model dumps generic filler questions and offloads its own reading | Questions **grouped by theme**, each shipping a **recommended default + one-line rationale** where a defensible technical default exists. Loop until no material ambiguity remains. |
| A pure-preference fork (names, aesthetics, tone) given a "default" | Inventing a preference falsely anchors the user | Present the options, invent **no** default - preference is the user's to set. |
| Emotion roles: "act as a hostile interrogator / a genius" | Emotion is a costume with no success criterion | **Job roles**: "a skeptical spec reviewer whose job is to find unhandled states." Job framing carries a pass/fail bar. |
| Model can offload work back onto the user | It asks instead of reading | **Anti-laziness valve**: "If it's answerable from the source, answer from the source and show your work." |
| "Be thorough / production-quality / fast" with no target | Invites quality/performance theater | **State the honest constraint** the model would hand-wave - especially when the real bottleneck lives outside the code being changed. |
| Output that must outlive the chat, left in chat | Chat evaporates; "why did we do X?" is unanswerable in a month | **Artifact contract**: name deliverable files suited to the task (SPEC/PLAN fit a rewrite; a review prompt may need none), record decisions with rationale, make milestones runnable + verifiable. Never inject document deliverables into a task that produces no documents. |
| Names an interactive feature the runtime may not have | Prompt targets an abstract chat, not the real harness | **Match the harness**: target facilities the current one actually offers (verify first), else fall back to batched questions + wait. |
| Freeform output / no output contract | Model picks the lowest-effort shape - a "looks good" blob - so results aren't auditable and filler hides easily | **Fixed output structure**: name the sections or fields the response must contain (e.g. findings ranked by severity, decisions + rationale, a coverage line) so the shape can't collapse into filler. |
| An irreversible or high-stakes action gated on the model's own judgment | Model performs the action on an unverified guess; damage is silent and unrecoverable | **Reversibility/confidence gate**: require a confirmation step, a confidence threshold, or a dry-run/undo path before the action fires - never let a fallible judgment trigger an unrecoverable effect unchecked. |
## Worked micro-example
**Input row:** "Ask the user as many clarifying questions as possible before starting."
**Diagnose:** model will generate a long list of shallow questions, many answerable from
the spec, to look thorough. Optimizes count. No stop condition.
**Rewrite:** "Read SPEC.md. For each genuine fork you find, ask one question grouped with
its siblings by theme, each with your recommended default and a one-line why. Anything
answerable from the spec, answer yourself and cite the line. Stop when no material
ambiguity remains; if you hit 3 rounds, list any leftovers as assumptions and proceed."
**Change table row:** `"as many questions as possible" -> shallow filler optimizing count
-> themed batches + defaults + anti-laziness valve + stop condition + round cap`.
## Saving the result (optional)
By default the refinement is delivered to chat only. When the user asks to save it
("save this", "write it to a file", "save to <path>"), also write the result to a
markdown file - this applies whether the run used the full method or a quick tweak.
- **Where.** Default to `.claude/prompts/` in the current project (create the directory if
missing). If the user names a path or directory, use that instead.
- **Filename.** `YYYY-MM-DD-<slug>.md`, where `<slug>` is a short kebab-case name derived
from the prompt's purpose (e.g. `2026-07-27-triage-agent-prompt.md`). If that file already
exists, append `-2`, `-3`, etc. - never overwrite.
- **Contents.** A single self-contained markdown document:
# <title from the prompt's purpose>
Refined <date> with refine-prompt.
## Original prompt
<the prompt as received, verbatim>
## Refined prompt
<the refined prompt>
## Changes
<the 3-column change table from Verify>
The file includes the original prompt even though the chat deliverable does not, so the
before/after stays auditable once the chat is gone.
- After writing, report the path. Still show the deliverable in chat unless the user asked
to save only. On a null verdict there is nothing to compile - say so and write no file.
## Reference
`references/principles.md` - the rationale behind each mechanism, and which ones carry
the most weight. Read it when a mechanism's fit is unclear or you are extending the
smell table.