Skills DirectorySkills Directory
SkillsLearnSecurityCategoriesDocsCommunityBlog
Sign InSubmit Skill
Skills Directory

Security-tested agent skills for Claude, coding agents, and AI workflows.

Directory

  • Browse Skills
  • All Skills A–Z
  • Claude Skills
  • Claude Code Skills
  • Agent Skills
  • Categories
  • Submit a Skill

Learn

  • Learn Hub
  • Install Claude Skills
  • Write SKILL.md
  • Skills vs MCP
  • Directories Compared

Security

  • Security
  • Methodology
  • Secure Claude Skills
  • Security Badges

Company

  • About
  • Community
  • Blog
  • API Docs
  • Advertise

2026 Skills Directory. All rights reserved.

Back to skills

Agent Evaluation

ASecurity

Designs and runs reproducible evaluations for AI agents, prompts, tools, skills, and model-backed workflows using realistic datasets, isolated baselines, objective assertions, rubric grading, trajectory analysis, cost/latency tracking, and regression comparison. Use when measuring agent quality, optimizing skill triggering, comparing prompts or models, or gating an AI feature release. Not for ordinary deterministic unit tests.

95 stars
0 votes
0 copies
0 views
Added 9/22/2026
ai-agents

Security Analysis

A100/100

Scanned 9/22/2026

Install to Claude Code

$npx -y skills add thiientv/godmode --skill agent-evaluation --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Agent Evaluation?

Add the live security badge to your README — it updates automatically with every re-scan.

Security grade badge for Agent Evaluation
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/thiientv-agent-evaluation/badge)](https://www.skillsdirectory.com/skills/thiientv-agent-evaluation)

More formats (shields.io, HTML) on the badges page.

Download Zip
Files
SKILL.md
---
name: agent-evaluation
description: >-
  Designs and runs reproducible evaluations for AI agents, prompts, tools,
  skills, and model-backed workflows using realistic datasets, isolated
  baselines, objective assertions, rubric grading, trajectory analysis,
  cost/latency tracking, and regression comparison. Use when measuring agent
  quality, optimizing skill triggering, comparing prompts or models, or gating
  an AI feature release. Not for ordinary deterministic unit tests.
---

# Agent Evaluation

Build a quality flywheel that can distinguish a real improvement from a lucky
run.

## Define the evaluation contract

Name the target behavior, users, risks, baseline, candidate, environment,
stochastic settings, and decision threshold. Start with a few realistic cases,
including a boundary or failure case. Split trigger-query optimization into a
fixed training set and held-out validation set.

Use [eval-schema.md](references/eval-schema.md) for cases, assertions, timing,
and result records.

## Run isolated comparisons

1. Snapshot the baseline before changing the candidate.
2. Run baseline and candidate on identical inputs in fresh contexts with no
   leaked expected answer or previous trace.
3. Capture final artifacts, public transcript/tool summaries, duration, token
   or request cost, and failures.
4. Grade deterministic assertions first; use a blinded rubric or human review
   for qualities that cannot be measured mechanically.
5. Repeat stochastic cases enough to expose variance. Do not hide flakiness by
   dropping inconvenient runs.

Measure task success, instruction adherence, tool selection and arguments,
trajectory efficiency, grounding, safety, output quality, latency, and cost only
when relevant. A single aggregate score must not hide a release-blocking metric.

## Analyze and iterate

Cluster repeated failures by cause, change one owning layer, rerun the affected
cases, then run the regression set. Compare candidate against baseline and
reject improvements that regress a protected metric beyond its tolerance.

Use `writing-skills` for skill-specific authoring and `release-engineering` for
production promotion. Never claim a score that was not read from an actual
result artifact.

## Completion condition

Cases, environment, baseline, candidate, artifacts, graders, costs, and limits
are reproducible; the decision follows predefined thresholds rather than a
post-hoc interpretation of the preferred result.

Attribution

thiientvthiientv
View sourceMore from thiientv →
SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments (0)

No comments yet. Be the first to comment!

SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Related Skills

Caveman

Ultra-compressed communication mode that cuts output tokens while keeping technical accuracy. Levels: lite, full, ultra and the wenyan variants. Use for /caveman, "caveman mode", "talk like caveman", "be brief" or "less tokens".

1066601 votes

Hyperplan

Adversarial multi-agent planning skill. Self-orchestrates 5 hostile category members (unspecified-low, unspecified-high, deep, ultrabrain, artistry) via team-mode for ruthless cross-critique debate, distills only the defensible insights, then MANDATORILY hands the distilled insight bundle to the `plan` agent for executable plan formalization. Use when planning needs maximum rigor and surfacing of weak assumptions, blind spots, and over-engineering. Triggers: 'hyperplan', 'hpp', '/hyperplan', ...

686011 votes

Mcp Code Execution

Routes multi-tool workflows through MCP servers for large datasets and pipelines. Use when Bash tool overhead is limiting throughput on data-heavy tasks.

3351 votes

catchup

Recovers the conversation and failed tool calls of a previous Codex, Claude Code, Antigravity, Cline, Copilot CLI, Cursor, DeepSeek Harness, Kimi, OpenCode, Pi Agent, or ZCode session. Use when the user says "catch up", "what did the last session do", "get me up to speed", "I switched agents", asks to recover/summarize a previous session before continuing, or asks to diagnose or report a catchup failure. Do NOT use for the current conversation, git history, or any non-agent log.

651 votes

math-skill

A comprehensive mathematical reasoning skill for AI assistants — handles arithmetic to research-level problems with rigorous step-by-step reasoning, systematic verification, and transparent uncertainty handling

381 votes
View all in ai-agents →