Skills DirectorySkills Directory
SkillsLearnSecurityCategoriesDocsBlogPro
Sign InSubmit Skill
Skills Directory

Security-tested agent skills for Claude, coding agents, and AI workflows.

Directory

  • Browse Skills
  • All Skills A–Z
  • Claude Skills
  • Claude Code Skills
  • Agent Skills
  • Categories
  • Authors
  • Submit a Skill

Learn

  • Learn Hub
  • Install Claude Skills
  • Write SKILL.md
  • Skills vs MCP
  • Directories Compared

Security

  • Security
  • Methodology
  • Secure Claude Skills
  • Security Badges
  • Chrome Extension
  • Skill Manager

Company

  • About
  • Community
  • Blog
  • API Docs
  • Advertise

2026 Skills Directory. All rights reserved.

ProTermsPrivacyRefunds
Back to skills

Evaluate And Harden Agent

ASecurity

Evaluate an agent and harden its failure modes before it touches real traffic — an offline eval harness scored on task-completion, trajectory, and tool-use correctness against a fixed task set, plus loop hardening (step/tool-call caps, timeouts + retries, stop conditions, human-in-the-loop on irreversible actions) and tracing so every step/tool-call/token-cost is observable, with cost and latency reported alongside quality. Reach for this when the user asks 'how do I know this agent works?', ...

7 stars
0 votes
0 copies
0 views
Added 9/23/2026
ai-agentsgorailsapi

Works with

api

Security Analysis

A100/100

Scanned 9/23/2026

$npx -y skills add mcorbett51090/RavenClaude --skill evaluate-and-harden-agent --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Evaluate And Harden Agent?

Add the live security badge to your README — it updates automatically with every re-scan.

Security grade badge for Evaluate And Harden Agent
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/mcorbett51090-evaluate-and-harden-agent/badge)](https://www.skillsdirectory.com/skills/mcorbett51090-evaluate-and-harden-agent)

More formats (shields.io, HTML) on the badges page. Keep it an A: scan every change in CI with Pro.

Download with Pro
Files
SKILL.md
---
name: evaluate-and-harden-agent
description: "Evaluate an agent and harden its failure modes before it touches real traffic — an offline eval harness scored on task-completion, trajectory, and tool-use correctness against a fixed task set, plus loop hardening (step/tool-call caps, timeouts + retries, stop conditions, human-in-the-loop on irreversible actions) and tracing so every step/tool-call/token-cost is observable, with cost and latency reported alongside quality. Reach for this when the user asks 'how do I know this agent works?', 'set up agent evals', or 'the agent loops / does something dangerous'. Used by `agent-implementation-engineer` (primary)."
---

# Skill: evaluate-and-harden-agent

> **Invoked by:** `agent-implementation-engineer` (primary). Also consulted by `agentic-systems-architect` when a go/no-go needs an evidence bar defined before build.
>
> **When to invoke:** "how do I know this agent actually works?"; "set up evals for the agent"; "the agent loops forever / burns tokens"; "the agent did something irreversible it shouldn't have"; "it works in the demo but fails in the tail".
>
> **Output:** an offline agent-eval harness (task-completion + trajectory + tool-use, over a fixed task set) and a hardening plan (caps, timeouts/retries, stop conditions, human-in-the-loop, tracing) — with cost and latency reported next to quality.

## Procedure

1. **Build a fixed task set before scoring anything.** Collect representative tasks with known-good outcomes (start with real or realistic inputs, include the hard/tail cases, not just the happy path). This set is the regression harness — freeze it and grow it as failures surface.
2. **Score three dimensions, not one.**
   - **Task-completion:** did the agent achieve the goal? (the outcome, judged against the known-good result — exact-match, rubric, or LLM-judge as fits.)
   - **Trajectory:** did it take a *sane path*, or flail and get lucky? Check step count, redundant/looping calls, and whether the plan was coherent — a right answer via a 40-step random walk is a latent failure.
   - **Tool-use correctness:** right tools, right arguments, recovered from errors? Wrong-tool and bad-argument rates are the leading indicators of agent quality.
3. **Run it offline against fixtures/mocks first.** Mock the tools/systems so the eval is deterministic, free, and fast — catch regressions before spending on live calls. Only after offline is green do you evaluate against live dependencies.
4. **Report cost and latency alongside quality — always.** Per-task **token cost** and **wall-clock**, plus the distribution (the p95, not just the mean). A correct agent that costs $2 and takes 90s per run may be unshippable; the number needs a price tag.
5. **Harden the loop against its failure modes.**
   - **Caps:** a hard **step/tool-call limit** so a stuck agent stops instead of billing forever.
   - **Timeouts + retries:** per-tool timeout with bounded backoff; a hung tool must not hang the agent.
   - **Stop conditions:** explicit goal-reached / no-progress / cap-hit exits.
   - **Human-in-the-loop:** a confirmation gate on **every irreversible or high-blast action** (send, pay, delete, write-to-prod) — the loop pauses for approval.
   - **Tracing:** every step, tool call, argument, result, and token cost recorded, so any failure is replayable.
6. **Feed production back into the eval set.** Treat early live runs as data — every new failure mode becomes a fixture in the frozen task set, so the agent can't regress on it again.

## Output format

- **Eval harness:** the task set + the three scorers (task-completion / trajectory / tool-use) + offline-fixture setup.
- **Quality + cost/latency report:** per-dimension scores with token-cost and wall-clock (mean + p95).
- **Hardening plan:** caps, timeouts/retries, stop conditions, human-in-the-loop gates, tracing wiring.
- **Volatile facts** (model IDs, pricing, framework tracing APIs): dated + `[verify-at-use]`.

## Guardrails

- **A demo is not an eval** — score the tail on a frozen task set, not the happy path once.
- **A right answer via a bad trajectory is a latent failure** — always score the path, not just the outcome.
- **Never ship an agent with a real-world write and no human gate** on the irreversible step.
- **Never report quality without cost and latency** next to it.

Attribution

mcorbett51090mcorbett51090
View sourceSee grades on GitHubMore from mcorbett51090 →
SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments (0)

No comments yet. Be the first to comment!

SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Related Skills

Caveman

Terse caveman voice: answer first, fluff gone, every technical fact kept. Use for /caveman, "caveman mode", "talk like caveman", "be brief", "less tokens". Stays on until "stop caveman" or "normal mode".

1100021 votes

Hyperplan

Adversarial multi-agent planning skill. Self-orchestrates 5 hostile category members (unspecified-low, unspecified-high, deep, ultrabrain, artistry) via team-mode for ruthless cross-critique debate, distills only the defensible insights, then MANDATORILY hands the distilled insight bundle to the `plan` agent for executable plan formalization. Use when planning needs maximum rigor and surfacing of weak assumptions, blind spots, and over-engineering. Triggers: 'hyperplan', 'hpp', '/hyperplan', ...

698461 votes

Writing Skills

Create and manage Claude Code skills in HASH repository following Anthropic best practices. Use when creating new skills, modifying skill-rules.json, understanding trigger patterns, working with hooks, debugging skill activation, or implementing progressive disclosure. Covers skill structure, YAML frontmatter, trigger types (keywords, intent patterns), UserPromptSubmit hook, and the 500-line rule. Includes validation and debugging with SKILL_DEBUG. Examples include rust-error-stack, cargo-dep...

3931 votes

Mcp Code Execution

Routes multi-tool workflows through MCP servers for large datasets and pipelines. Use when Bash tool overhead is limiting throughput on data-heavy tasks.

3421 votes

catchup

Recovers the conversation and failed tool calls of a previous Codex, Amp, Claude Code, Antigravity, Cline, Copilot CLI, Cursor, DeepSeek Harness, Grok Build, Kimi, OpenCode, Pi Agent, or ZCode session. Use when the user says "catch up", "what did the last session do", "get me up to speed", "I switched agents", asks to recover/summarize a previous session before continuing, or asks to diagnose or report a catchup failure. Do NOT use for the current conversation, git history, or any non-agent log.

741 votes
View all in ai-agents →