Skills DirectorySkills Directory
SkillsLearnSecurityCategoriesDocsBlogPro
Sign InSubmit Skill
Skills Directory

Security-tested agent skills for Claude, coding agents, and AI workflows.

Directory

  • Browse Skills
  • All Skills A–Z
  • Claude Skills
  • Claude Code Skills
  • Agent Skills
  • Categories
  • Authors
  • Submit a Skill

Learn

  • Learn Hub
  • Install Claude Skills
  • Write SKILL.md
  • Skills vs MCP
  • Directories Compared

Security

  • Security
  • Methodology
  • Secure Claude Skills
  • Security Badges
  • Chrome Extension
  • Skill Manager

Company

  • About
  • Community
  • Blog
  • API Docs
  • Advertise

2026 Skills Directory. All rights reserved.

ProTermsPrivacyRefunds
Back to skills

Ai Eval Plan

ASecurity

Plan an evaluation from examples the user already has, with a pass rule they can apply. Use when the user mentions AI eval, test the model, evaluation set, how will we know it works, or asks for a eval plan. AI in the business skill by Yasir Jilani.

2 stars
0 votes
0 copies
1 views
Added 9/30/2026
ai-agentspythongoawsgitapi

Works with

cliapi

Security Analysis

A100/100

Scanned 9/30/2026

$npx -y skills add SYasJ/claude-business-skills --skill ai-eval-plan --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Ai Eval Plan?

Add the live security badge to your README — it updates automatically with every re-scan.

Security grade badge for Ai Eval Plan
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/syasj-ai-eval-plan/badge)](https://www.skillsdirectory.com/skills/syasj-ai-eval-plan)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
Files
SKILL.md
---
name: ai-eval-plan
description: "Plan an evaluation from examples the user already has, with a pass rule they can apply. Use when the user mentions AI eval, test the model, evaluation set, how will we know it works, or asks for a eval plan. AI in the business skill by Yasir Jilani."
license: MIT
compatibility: Agent Skills standard. No network access, extra packages, or credentials required.
metadata:
  author: Yasir Jilani
  version: "1.0.0"
  domain: ai
---

<!-- GENERATED FILE - edits here are overwritten by scripts/generate.py.
     Edit the 'ai-eval-plan' entry in source/, then run:
       python3 scripts/generate.py && python3 scripts/validate.py
     See CONTRIBUTING.md. -->

# AI Eval Plan

Plan an evaluation from examples the user already has, with a pass rule they can apply.

## When to use this skill

Use this skill when the user:

- AI eval
- test the model
- evaluation set
- how will we know it works

## When not to use this skill

- The user wants a different domain's specialist skill.
- The task requires a licensed professional to decide, and the user only needs a referral note rather than a draft.
- The request asks you to deceive, evade a control, or hide material facts.

## Professional boundary

Do not write prompts or workflows that weaken safety rules, hide required disclosure, or invent model scores. This is not a certification and not a reason to send private data to a vendor.

## Operating boundaries

- Use only information the user provides or files they explicitly ask you to read. Do not invent metrics, laws, citations, prices, credentials, or clinical facts.
- Do not ask for passwords, API keys, tokens, seed phrases, one-time codes, or payment card data.
- Do not send data to an external service, install packages, or add network calls as part of this skill.
- Separate facts, assumptions, and recommendations. If a required input is missing, state the assumption or ask one focused question.
- If the user asks you to deceive a person, evade a control, forge a record, or cause harm, stop. Offer a legitimate alternative.
- Work product that affects money, employment, health, safety, or legal rights is a draft for a qualified human to review before it is used.

## Inputs to collect

- Ten or more real examples they can share
- What a pass looks like
- Who labels
- What they will not test yet

## Workflow


### 1. Step 1

Use their examples. Do not invent gold answers.
### 2. Step 2

Write the pass rule as something a reviewer can mark.
### 3. Step 3

Separate a smoke test from a full eval.
### 4. Step 4

Name who labels and how disagreements are kept.
### 5. Step 5

State what this plan does not prove.
### 6. Step 6

Do not publish a score from a plan that has not been run.

## Output

Deliver a **eval plan**.

- Purpose of this eval plan, in two sentences.
- Facts the user supplied, listed separately from assumptions.
- The work itself, in the structure the workflow names.
- Open questions, risks, and the single next action with an owner.
- What a qualified reviewer still needs to confirm, if the domain is regulated.

## Quality bar

- Every number, date, name, and citation came from the user or is marked as an assumption.
- The artifact can be used without reading this skill again.
- Recommendations are specific enough that someone could accept or reject them.
- Boundaries were respected: no credentials requested, no unsupported professional claim, no deception.

## Example

### Scenario

A slide says the support draft model is 95 percent accurate. Jonah has 12 real tickets and no labels. He needs a plan, not the slide.

### Example data

```text
examples on hand: tickets 4401-4412, exported 14 Sep 2026
labels: none
pass rule the lead will accept: no invented refund, no date that is not in the ticket
labeler: Rita, 2 hours on 22 Sep
will not test yet: tone score, legal advice, other languages
```

### Example outcome

**Eval plan**
The 95 percent comes off the slide. Nothing has been labeled.

Pass on each ticket: every dollar amount and date in the draft appears in the ticket. If not, fail.
Who labels: Rita, 22 September, tickets 4401-4412.
This is a 12-ticket smoke test. It does not prove a rate for the whole queue.
Not in this plan: tone, legal questions, languages other than the English in these tickets.
Next: Jonah does not quote a percent until Rita's marks are in a sheet.

## Anti-patterns

- Invented test cases presented as customer data
- A score with no labeled set
- A pass rule that is 'looks good'

## Related skills

- `ai-output-check`
- `prompt-revision`

Attribution

SYasJSYasJ
View sourceSee grades on GitHubMore from SYasJ →
SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments (0)

No comments yet. Be the first to comment!

SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Related Skills

Caveman

Terse caveman voice: answer first, fluff gone, every technical fact kept. Use for /caveman, "caveman mode", "talk like caveman", "be brief", "less tokens". Stays on until "stop caveman" or "normal mode".

1100021 votes

Hyperplan

Adversarial multi-agent planning skill. Self-orchestrates 5 hostile category members (unspecified-low, unspecified-high, deep, ultrabrain, artistry) via team-mode for ruthless cross-critique debate, distills only the defensible insights, then MANDATORILY hands the distilled insight bundle to the `plan` agent for executable plan formalization. Use when planning needs maximum rigor and surfacing of weak assumptions, blind spots, and over-engineering. Triggers: 'hyperplan', 'hpp', '/hyperplan', ...

698431 votes

Writing Skills

Create and manage Claude Code skills in HASH repository following Anthropic best practices. Use when creating new skills, modifying skill-rules.json, understanding trigger patterns, working with hooks, debugging skill activation, or implementing progressive disclosure. Covers skill structure, YAML frontmatter, trigger types (keywords, intent patterns), UserPromptSubmit hook, and the 500-line rule. Includes validation and debugging with SKILL_DEBUG. Examples include rust-error-stack, cargo-dep...

3931 votes

Mcp Code Execution

Routes multi-tool workflows through MCP servers for large datasets and pipelines. Use when Bash tool overhead is limiting throughput on data-heavy tasks.

3421 votes

catchup

Recovers the conversation and failed tool calls of a previous Codex, Amp, Claude Code, Antigravity, Cline, Copilot CLI, Cursor, DeepSeek Harness, Grok Build, Kimi, OpenCode, Pi Agent, or ZCode session. Use when the user says "catch up", "what did the last session do", "get me up to speed", "I switched agents", asks to recover/summarize a previous session before continuing, or asks to diagnose or report a catchup failure. Do NOT use for the current conversation, git history, or any non-agent log.

741 votes
View all in ai-agents →