Skip to content
Back to skills

Refine Prompt

ASecurity

Use when writing, refining, or reviewing any prompt or agent instruction that drives real work - agent/subagent prompts, spec/plan/rewrite prompts, interview or interrogation prompts (including designing the questions an agent should ask), meta-prompts, system prompts, and other model-facing instruction payloads such as tool descriptions, eval-judge rubrics, and grading criteria. Trigger whenever the user says "improve this prompt", "write a prompt for X", "refine my instructions", "why does ...

  • 2 stars
  • 0 votes
  • 0 copies
  • 0 views
  • Added October 9, 2026
ai-agentsrustgoshellperformance

Works with

  • cli

Security analysis

A100/100

Pro scans all 4 files and shows the line behind each finding

Scanned October 9, 2026

npx -y skills add zyeri/claude-refine-prompt --skill refine-prompt --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Refine Prompt?

Add the live security badge to your README. It updates with every re-scan.

Security grade badge for Refine Prompt
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/zyeri-refine-prompt/badge)](https://www.skillsdirectory.com/skills/zyeri-refine-prompt)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
SKILL.md
---
name: refine-prompt
description: >-
  Use when writing, refining, or reviewing any prompt or agent
  instruction that drives real work - agent/subagent prompts, spec/plan/rewrite prompts,
  interview or interrogation prompts (including designing the questions an agent should
  ask), meta-prompts, system prompts, and other model-facing instruction payloads such as
  tool descriptions, eval-judge rubrics, and grading criteria. Trigger whenever the user
  says "improve this prompt", "write a prompt for X", "refine
  my instructions", "why does the model ignore this instruction", "the agent keeps picking
  the wrong tool", or hands over a prompt/spec to tighten - even if they never say the
  word "prompt". The point is to replace vague exhortations ("be thorough", "ask lots of
  questions") with mechanisms a model will actually execute. NOT for: shell command
  prompts, permission dialogs, UI copy/microcopy, UserPromptSubmit hooks, commit messages,
  prose or document editing, human-facing interview or survey design, or merely
  translating an existing prompt.
---

# Refine Prompt

A prompt is a program written for a model. Refining one is not editing for style -
it is predicting where the draft will make the model do something generic or lazy,
and replacing that spot with a mechanism that forces the behavior you actually want.
The governing idea: **vague words optimize for the wrong thing.** "Ask many questions"
optimizes question _count_; "cover every fork" optimizes _coverage_. Same intent, and
only one of them is executable.

## Scope gate - pick the depth before you start

| Situation                                                  | What to do                                                                                                              |
| ---------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------- |
| Prompt drives substantial work AND user wants it sharpened | **Full method** (Gather -> Diagnose -> Rewrite -> Verify)                                                               |
| Small ask / quick tweak                                    | **Inline substitutions only** - apply the smell table, skip artifact reading, batched interrogation, and deliverables   |
| Diagnose flags nothing - no row predicts filler            | **Null verdict** - one line saying no material change is needed, naming what you checked. Do not manufacture a rewrite. |
| Shell prompt, permission prompt, UI dialog, UPS hook       | **Not this skill** - leave it alone                                                                                     |

The scope gate matters because the full method has real cost (reading artifacts,
building a simulation table). Spending it on a one-line tweak is the same mistake as
"be thorough" - effort pointed at the wrong thing.

## The method: Gather -> Diagnose -> Rewrite -> Verify

### 1. Gather

- **Read the referenced artifact first** - when one exists, reading is proportionate,
  and the user hasn't declined. Extract the real fork points, latent bugs, and honest
  constraints, then _enumerate them by name_ in the prompt. A named decision can't be
  silently skipped; "leave no stone unturned" is an exhortation, "resolve these 7
  decisions" hands over the actual stones. No artifact / too large / user declined ->
  enumerate assumptions in the prompt instead.
- **Front-load context.** Put a one-paragraph summary of the target artifact inside the
  prompt itself, so it survives being pasted into a fresh session and the model orients
  before its first tool call.
- **Preserve the user's intent skeleton.** Note every goal in the original. The refined
  prompt is the original _compiled_, never replaced - each goal must map to a mechanism,
  and no new scope gets invented.

### 2. Diagnose - simulate, don't judge

Build a table, **one row per thing the model must act on** - merge rows sharing a fix,
cap it at **12**, and note any remainder in one line:

| Instruction (as written) | What the model will literally do | Verdict |
| ------------------------ | -------------------------------- | ------- |

Fill the middle column by role-playing a literal, slightly-lazy model - not the ideal
reader you hope for. If a row predicts generic behavior (filler clarifying questions,
restating the goal back, hedged "it depends" non-answers, offloading reading onto the
user), mark it for rewrite. You are cataloguing _predicted failures_, not stylistic
preferences - this is what gives the Verify step something concrete to check.

Consult the **smell table** below to name each failure and its fix.

### 3. Rewrite - swap smells for mechanisms

Rewrite every flagged row using the matching mechanism from the smell table. Keep the
prompt lean: a mechanism that doesn't change predicted behavior is dead weight, cut it.

### 4. Verify - re-simulate the rewrite

Re-run the simulation table against the _refined_ prompt. Iterate until no row predicts
filler, **capped at 3 rounds**; at the cap, ship it and list the rows still predicting
filler as stated known-weak spots. A rewrite is an unverified hypothesis until you do
this. Then deliver:

1. The refined prompt.
2. A short **change table**: `original phrase -> predicted failure -> mechanism applied`.
   This makes the refinement itself checkable by the user, rather than a black-box "trust me".

If Diagnose flagged nothing, deliver the null verdict instead - an empty change table
means there was nothing to compile.

## Smell table - the core substitutions

The recurring failure patterns and what to replace them with. This is also the entire
toolkit for a **quick tweak** (scope gate row 2).

| Smell in the draft                                                      | Why it fails                                                                                                    | Mechanism to replace it with                                                                                                                                                                                                                                                  |
| ----------------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Quantity words: "many / all / thorough / as much as possible"           | Optimizes the count, not coverage; no termination = no done-state                                               | Themed batches **sized to the actual fork count** (2 forks -> batch of 2), an explicit **stop condition**, and a **round cap**. At the cap, record leftovers as stated assumptions.                                                                                           |
| "Ask the user questions" with no shape                                  | Model dumps generic filler questions and offloads its own reading                                               | Questions **grouped by theme**, each shipping a **recommended default + one-line rationale** where a defensible technical default exists. Loop until no material ambiguity remains.                                                                                           |
| A pure-preference fork (names, aesthetics, tone) given a "default"      | Inventing a preference falsely anchors the user                                                                 | Present the options, invent **no** default - preference is the user's to set.                                                                                                                                                                                                 |
| Emotion roles: "act as a hostile interrogator / a genius"               | Emotion is a costume with no success criterion                                                                  | **Job roles**: "a skeptical spec reviewer whose job is to find unhandled states." Job framing carries a pass/fail bar.                                                                                                                                                        |
| Model can offload work back onto the user                               | It asks instead of reading                                                                                      | **Anti-laziness valve**: "If it's answerable from the source, answer from the source and show your work."                                                                                                                                                                     |
| "Be thorough / production-quality / fast" with no target                | Invites quality/performance theater                                                                             | **State the honest constraint** the model would hand-wave - especially when the real bottleneck lives outside the code being changed.                                                                                                                                         |
| Output that must outlive the chat, left in chat                         | Chat evaporates; "why did we do X?" is unanswerable in a month                                                  | **Artifact contract**: name deliverable files suited to the task (SPEC/PLAN fit a rewrite; a review prompt may need none), record decisions with rationale, make milestones runnable + verifiable. Never inject document deliverables into a task that produces no documents. |
| Names an interactive feature the runtime may not have                   | Prompt targets an abstract chat, not the real harness                                                           | **Match the harness**: target facilities the current one actually offers (verify first), else fall back to batched questions + wait.                                                                                                                                          |
| Freeform output / no output contract                                    | Model picks the lowest-effort shape - a "looks good" blob - so results aren't auditable and filler hides easily | **Fixed output structure**: name the sections or fields the response must contain (e.g. findings ranked by severity, decisions + rationale, a coverage line) so the shape can't collapse into filler.                                                                         |
| An irreversible or high-stakes action gated on the model's own judgment | Model performs the action on an unverified guess; damage is silent and unrecoverable                            | **Reversibility/confidence gate**: require a confirmation step, a confidence threshold, or a dry-run/undo path before the action fires - never let a fallible judgment trigger an unrecoverable effect unchecked.                                                             |

## Worked micro-example

**Input row:** "Ask the user as many clarifying questions as possible before starting."

**Diagnose:** model will generate a long list of shallow questions, many answerable from
the spec, to look thorough. Optimizes count. No stop condition.

**Rewrite:** "Read SPEC.md. For each genuine fork you find, ask one question grouped with
its siblings by theme, each with your recommended default and a one-line why. Anything
answerable from the spec, answer yourself and cite the line. Stop when no material
ambiguity remains; if you hit 3 rounds, list any leftovers as assumptions and proceed."

**Change table row:** `"as many questions as possible" -> shallow filler optimizing count
-> themed batches + defaults + anti-laziness valve + stop condition + round cap`.

## Saving the result (optional)

By default the refinement is delivered to chat only. When the user asks to save it
("save this", "write it to a file", "save to <path>"), also write the result to a
markdown file - this applies whether the run used the full method or a quick tweak.

- **Where.** Default to `.claude/prompts/` in the current project (create the directory if
  missing). If the user names a path or directory, use that instead.
- **Filename.** `YYYY-MM-DD-<slug>.md`, where `<slug>` is a short kebab-case name derived
  from the prompt's purpose (e.g. `2026-07-27-triage-agent-prompt.md`). If that file already
  exists, append `-2`, `-3`, etc. - never overwrite.
- **Contents.** A single self-contained markdown document:

      # <title from the prompt's purpose>

      Refined <date> with refine-prompt.

      ## Original prompt

      <the prompt as received, verbatim>

      ## Refined prompt

      <the refined prompt>

      ## Changes

      <the 3-column change table from Verify>

  The file includes the original prompt even though the chat deliverable does not, so the
  before/after stays auditable once the chat is gone.
- After writing, report the path. Still show the deliverable in chat unless the user asked
  to save only. On a null verdict there is nothing to compile - say so and write no file.

## Reference

`references/principles.md` - the rationale behind each mechanism, and which ones carry
the most weight. Read it when a mechanism's fit is unclear or you are extending the
smell table.

Files in this skill

  • SKILL.md13.3 KB
  • evals/evals.json6 KB
  • evals/fixtures/SPEC.md687 B
  • references/principles.md3.9 KB

Attribution

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments

Loading comments…