Skills DirectorySkills Directory
SkillsLearnSecurityCategoriesDocsCommunityBlog
Sign InSubmit Skill
Skills Directory

Security-tested agent skills for Claude, coding agents, and AI workflows.

Directory

  • Browse Skills
  • All Skills A–Z
  • Claude Skills
  • Claude Code Skills
  • Agent Skills
  • Categories
  • Authors
  • Submit a Skill

Learn

  • Learn Hub
  • Install Claude Skills
  • Write SKILL.md
  • Skills vs MCP
  • Directories Compared

Security

  • Security
  • Methodology
  • Secure Claude Skills
  • Security Badges

Company

  • About
  • Community
  • Blog
  • API Docs
  • Advertise

2026 Skills Directory. All rights reserved.

ProTermsPrivacyRefunds
Back to skills

Delegation Eval Loop

ASecurity

Generator-evaluator loop for delegated sub-agent work: parent evaluates, child generates. Use layered verification (Deterministic, Invariant Gate, Trajectory/reliability, Semantic), sprint contracts with checkbox acceptance criteria, and targeted re-steers via agent__messageToSession. Use when briefing with strict acceptance criteria, evaluating child results from agent__checkSession, preventing premature "done" without proof, or when delegate routes here for high-stakes handoffs. Triggers: g...

10 stars
0 votes
0 copies
0 views
Added 9/28/2026
ai-agentsgogit

Works with

terminal

Security Analysis

A100/100

Scanned 9/28/2026

Install to Claude Code

$npx -y skills add fritzprix/libr-agent --skill delegation-eval-loop --agent claude-code

Installs into .claude/skills of the current project.

Are you the author of Delegation Eval Loop?

Add the live security badge to your README — it updates automatically with every re-scan.

Security grade badge for Delegation Eval Loop
[![Security: A — Skills Directory](https://www.skillsdirectory.com/api/skills/fritzprix-delegation-eval-loop/badge)](https://www.skillsdirectory.com/skills/fritzprix-delegation-eval-loop)

More formats (shields.io, HTML) on the badges page.

Files
SKILL.md
---
name: delegation-eval-loop
description: >
  Generator-evaluator loop for delegated sub-agent work: parent evaluates,
  child generates. Use layered verification (Deterministic, Invariant Gate,
  Trajectory/reliability, Semantic), sprint contracts with checkbox acceptance
  criteria, and targeted re-steers via agent__messageToSession. Use when
  briefing with strict acceptance criteria, evaluating child results from
  agent__checkSession, preventing premature "done" without proof, or when
  delegate routes here for high-stakes handoffs. Triggers: generator-evaluator,
  generator evaluator, delegation eval, subagent evaluation, verify subagent,
  delegation loop, eval loop, proof before done, acceptance criteria,
  delegation-eval-loop.
---

# Delegation Eval Loop

**Generator–evaluator:** the child session **generates** (implements, researches, edits); the **parent** session **evaluates** and alone decides acceptance. Never treat the child's self-reported success as the final grade.

For spawn, isolation, workspace inheritance, and tool naming, follow **`delegate` first**. This skill owns briefing contracts, layered grading, reject/re-steer, and bounded retries.

## Core Rules

1. **Evidence over assertion**: "all tests pass" from the child is a hypothesis. Run or inspect deterministic verification in the parent (or from parent-visible raw output).
2. **Layered grading order**: Deterministic → Invariant → Trajectory/reliability → Semantic. Fail fast on earlier layers.
3. **Sprint contract before spawn**: Negotiate objective, authorized paths, and checkbox acceptance criteria in the brief. Do not invent pass criteria after the child returns.
4. **Targeted feedback**: On reject, send exact failure output via `agent__messageToSession` — not "try again".
5. **Bounded loop**: Cap at 3–5 cycles. Then escalate to the user with diagnostics; do not spin forever.
6. **Reliability hard-fail**: Incomplete, cancelled, timed-out, circuit-broken, or evidence-free runs fail evaluation even if the prose claims completion.

## The 4-Layer Evaluation Stack

| Layer | Check | Verification Method | Action on Failure |
|---|---|---|---|
| **1. Deterministic** | Build, types, tests, lint | Parent runs the agreed commands (or requires raw exit-0 output in evidence) | Reject with compiler/test log |
| **2. Invariant Gate** | Path/policy boundaries | `git status` / `git diff --stat` vs authorized paths | Reject; force revert of forbidden edits |
| **3. Trajectory / reliability** | Real execution, no fake done | Result text + actions: missing cmds, loops, cancel/incomplete signals | Reject; demand missing evidence or stop |
| **4. Semantic** | User intent, edge cases | Parent LLM vs original requirements (only after 1–3 pass) | Reject with specific missed requirements |

Composite cheerleading is forbidden: Layer 1 pass does **not** override Layer 2/3 failure.

## Workflow

### 1. Brief with a Sprint Contract

When calling `agent__spawnSession` (or assigning via `agent__messageToSession`), include:

- **Exact Objective** — concrete deliverable
- **Authorized paths / FORBIDDEN paths** — invariant gate inputs
- **Acceptance Criteria** — checkbox list of commands/checks that must pass (`- [ ]`)
- **Deliverable Channel** — status + raw evidence in **final text** (not child scratchpad)

See [handoff-templates.md](references/handoff-templates.md).

### 2. Poll and Capture Child Result

- `agent__checkSession` until terminal idle/error (or use `wait=true` when appropriate)
- Capture final text. If the session errored, cancelled, or returned empty/incomplete diagnostics → **Layer 3 fail**; do not soft-accept.

### 3. Execute Layered Evaluation

Do **not** present the result to the user until layers pass (or you escalate).

1. Deterministic checks in the child's effective workspace (see Metadata `workspace:`)
2. Invariant gates (`git status` / scope)
3. Trajectory & reliability (claims vs evidence; incomplete/circuit-break)
4. Semantic review only if 1–3 pass

Details: [evaluation-protocol.md](references/evaluation-protocol.md).

### 4. Reject & Re-steer

If any layer fails:

1. Increment iteration counter (abort if over max, typically 4)
2. Structured correction: which layer, exact error snippet, single next action
3. `agent__messageToSession(sessionId, message)`
4. Return to Step 2

Templates: [handoff-templates.md](references/handoff-templates.md).

### 5. Final Acceptance

When all layers pass, synthesize for the user with verified evidence (commands run, scope clean). Mark which criteria were checked.

Attribution

fritzprixfritzprix
View sourceMore from fritzprix →
SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Is this your skill, or is something wrong with this listing? Request removal or report an issue. Author removals are honored within 72 hours.

Comments (0)

No comments yet. Be the first to comment!

SSkills DirectorySkills Directory

Ship a skill? Prove it's safe.

Free 120-pattern security scan, letter grade, and an embeddable README badge.

Submit a skill

Related Skills

Caveman

Ultra-compressed communication mode that cuts output tokens while keeping technical accuracy. Levels: lite, full, ultra and the wenyan variants. Use for /caveman, "caveman mode", "talk like caveman", "be brief" or "less tokens".

1074701 votes

Hyperplan

Adversarial multi-agent planning skill. Self-orchestrates 5 hostile category members (unspecified-low, unspecified-high, deep, ultrabrain, artistry) via team-mode for ruthless cross-critique debate, distills only the defensible insights, then MANDATORILY hands the distilled insight bundle to the `plan` agent for executable plan formalization. Use when planning needs maximum rigor and surfacing of weak assumptions, blind spots, and over-engineering. Triggers: 'hyperplan', 'hpp', '/hyperplan', ...

695601 votes

Mcp Code Execution

Routes multi-tool workflows through MCP servers for large datasets and pipelines. Use when Bash tool overhead is limiting throughput on data-heavy tasks.

3351 votes

catchup

Recovers the conversation and failed tool calls of a previous Codex, Claude Code, Antigravity, Cline, Copilot CLI, Cursor, DeepSeek Harness, Kimi, OpenCode, Pi Agent, or ZCode session. Use when the user says "catch up", "what did the last session do", "get me up to speed", "I switched agents", asks to recover/summarize a previous session before continuing, or asks to diagnose or report a catchup failure. Do NOT use for the current conversation, git history, or any non-agent log.

691 votes

math-skill

A comprehensive mathematical reasoning skill for AI assistants — handles arithmetic to research-level problems with rigorous step-by-step reasoning, systematic verification, and transparent uncertainty handling

381 votes
View all in ai-agents →